Home / Companies / Braintrust / Blog / Post Details
Content Deep Dive

Sandboxed agent evals with Harbor

Blog post from Braintrust

Post Details
Company
Date Published
Author
Braintrust Team
Word Count
555
Company Posts That Month
21
Language
English
Hacker News Points
-
Post removed?
No
Summary

Harbor, a Python framework from the terminal-bench team for evaluating sandboxed agents in isolated Docker containers, now offers a native Braintrust plugin that synchronizes evaluation results for easier comparison, sharing, and analysis. Previously, each Harbor run produced local job directories containing task environments, agent instructions, verifier results, and artifacts that had to be retained or manually shared to compare outcomes across runs. With the plugin, each job syncs managed task datasets, trial-level experiment rows, normalized verifier scores, and optional agent trajectories in Harbor’s ATIF format to a Braintrust project, while also preserving a local manifest of synced data. Users can inspect failed trials, trace agent decisions down to individual tool calls, chart separate verifier rewards, and keep setup, agent actions, and verification in a unified evaluation timeline. The integration can be enabled during a Harbor run through plugin flags or environment variables, keeps the Braintrust API key outside the sandbox container, and can backfill previously completed jobs without rerunning agents or creating duplicate records.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.