Sandboxed agent evals with Harbor
Blog post from Braintrust
Harbor, a Python framework from the terminal-bench team for evaluating sandboxed agents in isolated Docker containers, now offers a native Braintrust plugin that synchronizes evaluation results for easier comparison, sharing, and analysis. Previously, each Harbor run produced local job directories containing task environments, agent instructions, verifier results, and artifacts that had to be retained or manually shared to compare outcomes across runs. With the plugin, each job syncs managed task datasets, trial-level experiment rows, normalized verifier scores, and optional agent trajectories in Harbor’s ATIF format to a Braintrust project, while also preserving a local manifest of synced data. Users can inspect failed trials, trace agent decisions down to individual tool calls, chart separate verifier rewards, and keep setup, agent actions, and verification in a unified evaluation timeline. The integration can be enabled during a Harbor run through plugin flags or environment variables, keeps the Braintrust API key outside the sandbox container, and can backfill previously completed jobs without rerunning agents or creating duplicate records.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.