| GPT-6 Astra is now available in Loop.
Use it for pattern detection, eval creation, trace debugging, and other producti… |
@braintrust |
Company |
Original |
2026-09-08 |
478 |
9 |
2 |
0 |
0 |
0 |
| The Braintrust MCP server now exposes write tools.
Your coding agent can author prompts, scorers, and classifiers, con… |
@braintrust |
Company |
Original |
2026-09-08 |
288 |
3 |
0 |
0 |
0 |
2 |
| Your coding agent already knows how you work. It's set up with your repo, tools, and skills. Asking you to move part of… |
@braintrust |
Company |
Original |
2026-09-04 |
593 |
6 |
0 |
2 |
0 |
0 |
| Your team deserves more from your agent observability platform - a single, connected place for instrumentation, investi… |
@braintrust |
Company |
Original |
2026-09-03 |
35,712 |
52 |
11 |
12 |
9 |
23 |
| Use the Braintrust JavaScript/TypeScript SDK for tracing and evaling agents in any JS or TS project.
Includes integrat… |
@braintrust |
Company |
Original |
2026-09-02 |
411 |
3 |
0 |
0 |
0 |
0 |
| Cloudflare Agents now emits native OpenTelemetry traces, and you can route them straight to Braintrust.
Export spans v… |
@braintrust |
Company |
Original |
2026-09-01 |
1,257 |
9 |
4 |
0 |
0 |
4 |
| Imagine doing your job without ever looking anything up on the internet. That's an agent without web search.
Giving ag… |
@braintrust |
Company |
Original |
2026-08-31 |
16,097 |
42 |
8 |
5 |
5 |
28 |
| The Braintrust eval library has a repo of skills your coding agent can read, so you can easily build and run evals on y… |
@braintrust |
Company |
Original |
2026-08-28 |
1,956 |
24 |
2 |
2 |
1 |
12 |
| RT @daRubberDuckiee: used this skill last night to audit my eval analysis (confidence intervals, paired comparisons, et… |
@braintrust |
Company |
Repost |
2026-08-28 |
54 |
0 |
2 |
0 |
0 |
0 |
| If you want to reduce agent costs without trading away quality, the right unit of analysis is cost per resolved request… |
@braintrust |
Company |
Original |
2026-08-27 |
502 |
6 |
2 |
1 |
0 |
2 |
| Rex is an AI-native service for automating order-to-cash. From day one, their engineers made a deliberate investment to… |
@braintrust |
Company |
Original |
2026-08-26 |
441 |
5 |
2 |
0 |
0 |
1 |
| The entire internet spent the last 48 hours saying the same thing: I don't want your agent, I want my agent to use your… |
@braintrust |
Company |
Original |
2026-08-26 |
16,500 |
32 |
2 |
3 |
4 |
16 |
| For a long-horizon agent making hundreds of decisions per trajectory, scoring the final answer is not enough to know wh… |
@braintrust |
Company |
Original |
2026-08-25 |
748 |
9 |
0 |
1 |
1 |
2 |
| Some agents need a sandbox where they can do work, liking editing files, installing dependencies, and launching builds.… |
@braintrust |
Company |
Original |
2026-08-24 |
2,061 |
16 |
1 |
3 |
0 |
12 |
| What's new:
- Braintrust's Eval library, a resource of open source evals and 20+ skills, for running your own evals an… |
@braintrust |
Company |
Original |
2026-08-21 |
808 |
11 |
0 |
3 |
0 |
8 |
| RT @daRubberDuckiee: @braintrust sf event: devTools demo night ❤️ https://t.co/xWbWmB1PHB |
@braintrust |
Company |
Repost |
2026-08-21 |
59 |
0 |
1 |
0 |
0 |
0 |
| There are two ways to score a coding agent. By its output, or by its behavior.
Output scoring tells you if the agent … |
@braintrust |
Company |
Original |
2026-08-20 |
1,560 |
13 |
1 |
2 |
1 |
5 |
| RT @imabhiprasad: cooked up a @braintrust plugin for @harborframework. Harbor feels like exactly the right abstraction … |
@braintrust |
Company |
Repost |
2026-08-19 |
53 |
0 |
3 |
0 |
0 |
0 |
| Join @ankrgyl at @browserbase's Navigate conference for a panel on the state of agent infra.
Thursday, Sep 10 in San F… |
@braintrust |
Company |
Original |
2026-08-19 |
435 |
9 |
0 |
0 |
0 |
0 |
| RT @daRubberDuckiee: basics of evals explained by seals https://t.co/5Sk3EDDuHl |
@braintrust |
Company |
Repost |
2026-08-19 |
52 |
0 |
3 |
0 |
0 |
0 |
| RT @iz_hurley_: Evals are hard, but also really fun. They are increasingly the ways we get to understand the systems th… |
@braintrust |
Company |
Repost |
2026-08-18 |
40 |
0 |
4 |
0 |
0 |
0 |
| RT @daRubberDuckiee: The team cooked on this one 🙂 I'm the MOST excited about open-sourcing our eval skills repo
they … |
@braintrust |
Company |
Repost |
2026-08-18 |
40 |
0 |
1 |
0 |
0 |
0 |
| Evals are hard. Good eval research should be accessible.
At Braintrust, we have unique access to how leading AI teams… |
@braintrust |
Company |
Original |
2026-08-18 |
22,783 |
76 |
10 |
6 |
4 |
92 |
| OpenAI recently shipped three new models in the GPT 5.6 family.
So we evaled them all on the key building blocks of ag… |
@braintrust |
Company |
Original |
2026-08-17 |
2,161 |
12 |
1 |
2 |
1 |
3 |
| Trace OpenAI Codex sessions in Braintrust with the `trace-codex` plugin.
Add the plugin to your Codex setup to capture… |
@braintrust |
Company |
Original |
2026-08-13 |
811 |
6 |
0 |
0 |
0 |
3 |
| Kimi K3 and DeepSeek V4 Flash are now available as built-in models on Braintrust, joining GLM-5.2. Open models keep imp… |
@braintrust |
Company |
Original |
2026-08-12 |
1,005 |
11 |
0 |
1 |
0 |
1 |
| Eve's legal agents analyze thousands of documents, build medical chronologies, summarize eight-hour depositions, run le… |
@braintrust |
Company |
Original |
2026-08-11 |
1,222 |
11 |
2 |
2 |
1 |
7 |
| New in the experiments page:
- The analysis chart now includes options for All scores (avg) and Scale by axis, so you … |
@braintrust |
Company |
Original |
2026-08-10 |
659 |
6 |
0 |
1 |
0 |
0 |
| Recent research by @a1zhang found that training with an RLM harness can help models learn strategies that transfer to n… |
@braintrust |
Company |
Original |
2026-08-07 |
743 |
9 |
0 |
0 |
0 |
5 |
| Cloudflare's dashboard agent spans its entire developer platform, from deploying Workers to debugging production instan… |
@braintrust |
Company |
Original |
2026-08-07 |
8,164 |
26 |
4 |
2 |
1 |
12 |
| Trust the agents you build on your data
A workshop from Braintrust and @motherduck. Aug 12 at 10AM PT. Led by @darubbe… |
@braintrust |
Company |
Original |
2026-08-06 |
586 |
5 |
1 |
0 |
0 |
1 |
| If you build agents with @Cloudflare, you can send traces to Braintrust to see how these agents behave, eval them, and … |
@braintrust |
Company |
Original |
2026-08-05 |
2,049 |
5 |
1 |
0 |
0 |
5 |
| Eval foundations, from Braintrust.
Learn to run evals so you can ship quality agents.
- Write deterministic and LLM-a… |
@braintrust |
Company |
Original |
2026-08-04 |
551 |
6 |
0 |
3 |
0 |
1 |
| Most early-stage teams know they need evals and observability. Cost and limited resources push it down the list.
Brain… |
@braintrust |
Company |
Original |
2026-08-03 |
1,066 |
4 |
0 |
2 |
0 |
0 |
| Which design MCP server builds a better front-end page, Figma or Paper?
We ran an independent eval in Braintrust with… |
@braintrust |
Company |
Original |
2026-08-03 |
1,727 |
13 |
1 |
2 |
1 |
7 |
| What's new:
- Behavior specs, an open standard for defining and evaling how an agent should behave over an entire traj… |
@braintrust |
Company |
Original |
2026-07-31 |
802 |
5 |
0 |
2 |
0 |
1 |
| Pylon's customers use its support platform to serve their customers, so the stakes for quality are high.
Because they … |
@braintrust |
Company |
Original |
2026-07-30 |
587 |
5 |
1 |
0 |
0 |
0 |
| To build long-horizon agents you can trust, you have to supervise the process, not just the final output. Agents that m… |
@braintrust |
Company |
Original |
2026-07-29 |
7,920 |
39 |
3 |
1 |
4 |
31 |
| Both GLM-5.2 and Opus 4.8 models achieved 83 out of 100 correct answers in our perturbation test, but the cost efficien… |
@braintrust |
Company |
Original |
2026-07-28 |
1,008 |
8 |
0 |
2 |
0 |
0 |
| Traditional observability is reactive. Active observability works continuously in the background, surfacing patterns in… |
@braintrust |
Company |
Original |
2026-07-24 |
598 |
4 |
0 |
4 |
0 |
2 |
| Trace OpenCode sessions with the Braintrust OpenCode plugin.
Install `@braintrust/trace-opencode` to capture sessions,… |
@braintrust |
Company |
Original |
2026-07-23 |
405 |
4 |
0 |
4 |
0 |
1 |
| Box's early evals were manual, with responses marked in a spreadsheet or a Box Note.
With Braintrust, they built a pra… |
@braintrust |
Company |
Original |
2026-07-22 |
5,533 |
14 |
1 |
4 |
1 |
7 |
| There's a claim that open source models aren't as good as closed source models once the context gets really long becaus… |
@braintrust |
Company |
Original |
2026-07-21 |
1,753 |
7 |
0 |
0 |
1 |
2 |
| Both Figma and Paper have MCP servers that give a coding agent a suite of design tools and claim to produce better fron… |
@braintrust |
Company |
Original |
2026-07-20 |
52,756 |
88 |
10 |
6 |
12 |
56 |
| Swapping in a cheaper model rarely improves overall cost-efficiency. Token costs may drop, but failures increase, retri… |
@braintrust |
Company |
Original |
2026-07-20 |
552 |
2 |
0 |
1 |
0 |
0 |
| What's new:
- Extend log retention up to 180 days to keep your activity history queryable for longer
- GLM-5.2 availab… |
@braintrust |
Company |
Original |
2026-07-17 |
469 |
3 |
0 |
1 |
0 |
0 |
| Topics clusters every production trace to surface what agents are doing. But 100% coverage only works with a small mode… |
@braintrust |
Company |
Original |
2026-07-15 |
6,964 |
16 |
4 |
2 |
3 |
10 |
| We evaled the GPT-5.6 family, plus Anthropic's Fable, Opus 4.8, and Sonnet 5, on the key building blocks of agentic wor… |
@braintrust |
Company |
Original |
2026-07-10 |
1,899 |
12 |
0 |
2 |
1 |
7 |
| If you're building a voice agent, picking the right speech-to-text model is not obvious. Every provider claims to be ac… |
@braintrust |
Company |
Original |
2026-07-09 |
746 |
10 |
0 |
3 |
0 |
2 |
| The Braintrust Go SDK enables automatic instrumentation with no code changes using Orchestrion, or the use of manual mi… |
@braintrust |
Company |
Original |
2026-07-08 |
493 |
2 |
0 |
1 |
0 |
0 |
| Phrase search breaks when every word is common but the exact sequence is rare. At 100TB+ of agent traces, queries time … |
@braintrust |
Company |
Original |
2026-07-08 |
6,781 |
8 |
0 |
1 |
1 |
8 |
| Agent failures rarely show up in your test suite. Production creates unexpected issues that curated datasets simply can… |
@braintrust |
Company |
Original |
2026-07-07 |
462 |
5 |
0 |
1 |
0 |
0 |
| We evaled who will win the US vs Belgium match.
Last week, we ran 48 matchups through 6 configurations of @p0 research… |
@braintrust |
Company |
Original |
2026-07-06 |
506 |
8 |
0 |
1 |
0 |
0 |
| We built a Braintrust-native eval in collaboration with Baseten to test whether GLM-5.2 can preserve exact long-context… |
@braintrust |
Company |
Original |
2026-07-06 |
7,250 |
20 |
2 |
2 |
1 |
3 |
| We ran the 48 Group Stage World Cup matchups through six configurations of @p0 research agents and scored every output … |
@braintrust |
Company |
Original |
2026-07-02 |
4,057 |
18 |
4 |
1 |
1 |
1 |
| The answer to "which model is cheapest" depends entirely on whether you're asking about cost per task or cost per succe… |
@braintrust |
Company |
Original |
2026-07-02 |
636 |
2 |
1 |
4 |
0 |
1 |
| Game, set, match on the inaugural Agent Open.
Thanks to everyone who stopped by and grabbed a paddle.
And thanks to o… |
@braintrust |
Company |
Original |
2026-07-01 |
1,156 |
16 |
1 |
0 |
0 |
1 |
| Run the best OSS models in Braintrust, in collaboration with @Baseten.
Call GLM-5.2 natively, eval its quality, and ob… |
@braintrust |
Company |
Original |
2026-06-30 |
9,754 |
56 |
5 |
4 |
3 |
17 |
| The AI teams that ship quality agents put evals and observability in place early.
Braintrust for Startups gives early… |
@braintrust |
Company |
Original |
2026-06-30 |
10,889 |
15 |
5 |
1 |
1 |
6 |
| Reading traces one by one doesn't scale.
Topics automatically clusters production traces so you can identify patterns,… |
@braintrust |
Company |
Original |
2026-06-29 |
495 |
4 |
0 |
1 |
0 |
0 |
| Run a full chess eval without writing a single line of code using the Braintrust CLI.
- Take a CSV of chess puzzles a… |
@braintrust |
Company |
Original |
2026-06-29 |
968 |
13 |
3 |
3 |
1 |
3 |
| What's new:
- Smarter evals with role-based score visibility and contextual rubrics.
- Secure access to OpenAI, Anthro… |
@braintrust |
Company |
Original |
2026-06-26 |
1,099 |
11 |
1 |
2 |
0 |
1 |
| Stateful agents that do real work are worth investing in, but they're also more difficult to eval. The hard part isn't … |
@braintrust |
Company |
Original |
2026-06-26 |
6,107 |
14 |
0 |
1 |
1 |
14 |
| Cost-efficiency doesn't mean picking the cheapest model. Model choice, routing, retries, fallbacks, and escalation all … |
@braintrust |
Company |
Original |
2026-06-25 |
381 |
2 |
0 |
1 |
0 |
0 |
| We analyzed 1,781 real agent traces from @huggingface to understand what actually drives agent success across models, b… |
@braintrust |
Company |
Original |
2026-06-25 |
9,832 |
32 |
12 |
6 |
0 |
10 |
| As models develop increasingly nuanced differences in reasoning style, tool use, and context handling, the industry nee… |
@braintrust |
Company |
Original |
2026-06-24 |
428 |
4 |
0 |
1 |
0 |
0 |
| The Agent Open panel lineup is live.
On Jun 30th in San Francisco Braintrust and friends host a conversation with lead… |
@braintrust |
Company |
Original |
2026-06-22 |
406 |
5 |
0 |
1 |
0 |
0 |
| There have been six generations of AI agents:
- A simple prompt that asks a model a question.
- A fixed pipeline tha… |
@braintrust |
Company |
Original |
2026-06-22 |
500 |
11 |
2 |
1 |
0 |
2 |
| When you're building AI systems, you need to know what prompt your LLM received, what it returned, and how many tokens … |
@braintrust |
Company |
Original |
2026-06-19 |
620 |
7 |
0 |
0 |
0 |
2 |
| How do you make AI traces readable for non-engineers?
Custom trace views in Braintrust transform a raw trace into a f… |
@braintrust |
Company |
Original |
2026-06-18 |
287 |
1 |
0 |
1 |
0 |
0 |
| AI governance is entering a new phase. The EU AI Act enforces legal accountability for any company with EU customers, a… |
@braintrust |
Company |
Original |
2026-06-16 |
372 |
4 |
0 |
0 |
1 |
1 |
| Every AI team is building with different frameworks, different model providers, and different languages.
Braintrust is… |
@braintrust |
Company |
Original |
2026-06-16 |
233 |
2 |
1 |
1 |
0 |
0 |
| How does your team rank when it comes to shipping quality AI products?
Braintrust's AI quality assessment maps your cu… |
@braintrust |
Company |
Original |
2026-06-12 |
383 |
3 |
0 |
1 |
0 |
0 |
| Your agent has brain rot https://t.co/VLEA3nTylP |
@braintrust |
Company |
Original |
2026-06-11 |
7,457 |
14 |
0 |
0 |
1 |
1 |
| Build a full eval pipeline (dataset, prompt, scorer, experiment) using just the Braintrust CLI and skills.
In this vid… |
@braintrust |
Company |
Original |
2026-06-09 |
357 |
2 |
0 |
0 |
0 |
0 |
| The inaugural Agent Open is happening June 30 in SF.
Come talk about AI observability, then show off your skills at o… |
@braintrust |
Company |
Original |
2026-06-08 |
994 |
7 |
2 |
1 |
0 |
1 |
| What's new:
-Topics is now GA, with $249 in credits for Pro plans
-Multi-user human review, with averaged scores
-Work… |
@braintrust |
Company |
Original |
2026-06-05 |
498 |
6 |
0 |
1 |
0 |
0 |
| Raw agent traces can include millions of tokens across hundreds of spans. Too large for direct embedding, too irregular… |
@braintrust |
Company |
Original |
2026-06-04 |
253 |
2 |
0 |
2 |
0 |
0 |
| Topics is now GA on all plans.
Continuously find the patterns worth investigating across your production traffic. htt… |
@braintrust |
Company |
Original |
2026-06-01 |
7,201 |
30 |
4 |
3 |
1 |
13 |
| Loop can create and manage dataset snapshots, tag them with environments, and prompt you to save before making changes.… |
@braintrust |
Company |
Original |
2026-05-29 |
529 |
5 |
0 |
1 |
0 |
0 |
| Most traditional enterprises gave responsibility for AI to their ML team, but the model providers own the data pipeline… |
@braintrust |
Company |
Original |
2026-05-28 |
657 |
2 |
0 |
2 |
0 |
0 |
| Thanks to @Redpoint and congratulations to all the companies included on the 2026 InfraRed 100. https://t.co/338zwWQzny |
@braintrust |
Company |
Original |
2026-05-27 |
238 |
3 |
0 |
0 |
0 |
0 |
| Zero-code AI observability for Java applications. Attach the Braintrust Java agent at JVM startup to automatically trac… |
@braintrust |
Company |
Original |
2026-05-27 |
371 |
2 |
0 |
1 |
0 |
0 |
| Without validation of what good looks like, it's impossible to judge whether AI quality is improving or regressing.
H… |
@braintrust |
Company |
Original |
2026-05-26 |
353 |
4 |
1 |
2 |
0 |
0 |
| What's new:
- Scatterplots and snapshot views on the Topics page to visualize trace clusters
- Comparison grades for e… |
@braintrust |
Company |
Original |
2026-05-25 |
918 |
6 |
1 |
1 |
0 |
2 |
| Agent design has evolved through six distinct generations as models have grown smarter and more capable.
From simple p… |
@braintrust |
Company |
Original |
2026-05-22 |
390 |
3 |
1 |
3 |
0 |
2 |
| AI observability has shifted from the traditional pillars of metrics, logs, and traces to a new set of challenges: trac… |
@braintrust |
Company |
Original |
2026-05-21 |
373 |
3 |
0 |
1 |
0 |
2 |
| https://t.co/LjMBwfYc9Z |
@braintrust |
Company |
Original |
2026-05-21 |
3,598 |
22 |
2 |
1 |
1 |
25 |
| We tested whether "bash is all you need" for AI agents by building an eval harness that compared SQL, bash, and filesys… |
@braintrust |
Company |
Original |
2026-05-20 |
493 |
2 |
0 |
3 |
1 |
1 |
| Streamline dashboard management by copying charts and entire dashboard views across projects and organizations.
Expor… |
@braintrust |
Company |
Original |
2026-05-19 |
390 |
2 |
0 |
1 |
0 |
0 |
| Five hard-learned lessons from teams running thousands of evals daily:
- Good evals enable 24‑hour model swaps, feed o… |
@braintrust |
Company |
Original |
2026-05-18 |
861 |
9 |
1 |
2 |
0 |
5 |
| Single-turn evals can't tell you if your chatbot asked for the same information twice or kept customers in polite loops… |
@braintrust |
Company |
Original |
2026-05-15 |
334 |
5 |
0 |
2 |
0 |
1 |
| Running evals locally ties up your machine and makes it hard to collaborate with teammates. Braintrust's Sandboxes feat… |
@braintrust |
Company |
Original |
2026-05-15 |
440 |
5 |
0 |
2 |
0 |
1 |
| https://t.co/Kj9mm2SpwZ |
@braintrust |
Company |
Original |
2026-05-14 |
506 |
3 |
1 |
0 |
1 |
1 |
| Going from prototype to production is more challenging than ever. With AI products, teams need to manage multi-step age… |
@braintrust |
Company |
Original |
2026-05-14 |
300 |
4 |
0 |
2 |
0 |
0 |
| https://t.co/fnYLtQdLKh |
@braintrust |
Company |
Original |
2026-05-14 |
549 |
3 |
1 |
0 |
0 |
3 |
| We've identified a security incident involving unauthorized access to an internal AWS account. We've communicated with … |
@braintrust |
Company |
Original |
2026-05-06 |
4,157 |
5 |
0 |
2 |
2 |
1 |
| At @vercel, customers expect to build with the latest models immediately, so they ship support within hours of release.… |
@braintrust |
Company |
Original |
2026-05-05 |
524 |
2 |
0 |
1 |
0 |
1 |
| https://t.co/adgzyTIF31 |
@braintrust |
Company |
Original |
2026-05-05 |
643 |
4 |
0 |
0 |
0 |
1 |
| Encyclopedia Evalica
A resource from Braintrust compiling the most important things to know about evals. The terms to … |
@braintrust |
Company |
Original |
2026-05-04 |
6,267 |
13 |
2 |
1 |
4 |
9 |
| https://t.co/leC0EMYURn |
@braintrust |
Company |
Original |
2026-05-04 |
666 |
6 |
1 |
0 |
1 |
2 |
| An eval platform is more than just a test runner. Evals require shared definitions of "good," reliable data pipelines, … |
@braintrust |
Company |
Original |
2026-05-01 |
465 |
3 |
0 |
1 |
0 |
1 |
| https://t.co/wdspP3xyhd |
@braintrust |
Company |
Original |
2026-05-01 |
576 |
5 |
0 |
0 |
0 |
3 |
| For AI PMs, evals are the new PRD.
At @PLEDalliance Summit New York, Ameya Bhatawdekar discussed the new product deve… |
@braintrust |
Company |
Original |
2026-04-30 |
286 |
4 |
1 |
1 |
0 |
0 |
| https://t.co/ea31fEYrYt |
@braintrust |
Company |
Original |
2026-04-29 |
379 |
3 |
0 |
0 |
0 |
1 |
| Braintrust x @Nasdaq
Thank you to @wing_vc and congratulations to everyone on this year's Enterprise Tech 30. https:/… |
@braintrust |
Company |
Original |
2026-04-27 |
677 |
10 |
3 |
1 |
0 |
2 |
| https://t.co/GesMzfbZAe |
@braintrust |
Company |
Original |
2026-04-27 |
480 |
2 |
0 |
0 |
1 |
1 |
| The timeline view now shows token distribution and cache hit rates across your trace spans.
Scale visualizations by to… |
@braintrust |
Company |
Original |
2026-04-26 |
378 |
7 |
0 |
10 |
0 |
0 |
| The bt setup command now auto-detects your programming language and asks fewer questions during onboarding.
One comman… |
@braintrust |
Company |
Original |
2026-04-24 |
264 |
3 |
0 |
1 |
0 |
0 |
| https://t.co/H3liOrlQ4M |
@braintrust |
Company |
Original |
2026-04-23 |
217 |
2 |
0 |
0 |
0 |
2 |
| https://t.co/WuA1Qa1XfW |
@braintrust |
Company |
Original |
2026-04-23 |
634 |
6 |
1 |
0 |
0 |
0 |
| You ran an eval and got a score. What comes next?
Start by reviewing 5-10 examples from your results. For each output,… |
@braintrust |
Company |
Original |
2026-04-21 |
267 |
3 |
0 |
1 |
0 |
0 |
| Evals 101: a new course from Braintrust. Everything you need to know about evals, and how to do them yourself.
Module … |
@braintrust |
Company |
Original |
2026-04-20 |
37,026 |
73 |
1 |
5 |
2 |
200 |
| Sandboxes push your agent eval code once and run it from the playground on demand.
Iterate on complex agent workflows … |
@braintrust |
Company |
Original |
2026-04-16 |
336 |
2 |
0 |
2 |
0 |
0 |
| Topics automatically clusters your production logs to surface patterns like user complaints, feature requests, and erro… |
@braintrust |
Company |
Original |
2026-04-15 |
380 |
5 |
0 |
1 |
0 |
0 |
| The EU AI Act and ISO 42001 require real-time audit evidence of AI system behavior.
AI observability provides the logg… |
@braintrust |
Company |
Original |
2026-04-14 |
529 |
2 |
0 |
1 |
0 |
0 |
| Coding agents are good at evals. They can read structured output, form hypotheses, make targeted edits, and verify the … |
@braintrust |
Company |
Original |
2026-04-13 |
797 |
3 |
1 |
1 |
1 |
1 |
| When your AI makes decisions across dozens of steps, a single-turn scorer can’t catch where the conversation went wrong… |
@braintrust |
Company |
Original |
2026-04-11 |
327 |
1 |
0 |
1 |
0 |
0 |
| What's new:
-Sandboxes for agent evals: Push once, run from playground
-Trajectory scoring: Token visualization for LL… |
@braintrust |
Company |
Original |
2026-04-10 |
583 |
6 |
3 |
1 |
0 |
1 |
| Browserbase provides the infrastructure that lets AI browse the web.
When @browserbase customers combine internal mode… |
@braintrust |
Company |
Original |
2026-04-09 |
1,595 |
11 |
2 |
1 |
0 |
2 |
| “Shipping on vibes” is how AI breaks. Real evals ensure AI products actually work.
Hear more from @darubberduckiee at … |
@braintrust |
Company |
Original |
2026-04-09 |
574 |
4 |
0 |
0 |
0 |
0 |
| Braintrust in NYC.
Up now in the Broadway-Lafayette station. https://t.co/3jOJbsMNOx |
@braintrust |
Company |
Original |
2026-04-08 |
44,295 |
19 |
3 |
0 |
1 |
1 |
| Braintrust is going to @HumanXCo.
Join us for AI After Hours at SFMOMA with friends from @Browserbase, @Modal, @Resol… |
@braintrust |
Company |
Original |
2026-04-06 |
974 |
9 |
1 |
1 |
1 |
0 |
| Want to run an eval on a multi-turn conversation between a person and AI?
Instead of scoring each span independently, … |
@braintrust |
Company |
Original |
2026-04-03 |
323 |
2 |
0 |
0 |
0 |
1 |
| https://t.co/u8kP9hEH36 |
@braintrust |
Company |
Original |
2026-04-01 |
102,579 |
129 |
13 |
2 |
11 |
364 |
| Congratulations to @trybasis, @Browserbase, @meetgranola, @Lovable, @Modal, @Vercel and all the friends and customers o… |
@braintrust |
Company |
Original |
2026-03-31 |
310 |
4 |
1 |
0 |
0 |
0 |
| Thank you @tryramp. https://t.co/tP6gd7RVoc |
@braintrust |
Company |
Original |
2026-03-31 |
359 |
5 |
0 |
0 |
0 |
0 |
| The trace tree now shows estimated LLM costs inline on each span, with costs automatically propagated from child to par… |
@braintrust |
Company |
Original |
2026-03-31 |
224 |
2 |
0 |
1 |
0 |
1 |
| What's new:
-SQL sandbox linting: catch errors and slow queries before you run
-Costs in trace tree: see spend on each… |
@braintrust |
Company |
Original |
2026-03-27 |
444 |
3 |
1 |
1 |
0 |
2 |
| If experiments in Braintrust are like making a pull request to your repository, eval playgrounds are like editing code … |
@braintrust |
Company |
Original |
2026-03-26 |
244 |
2 |
0 |
1 |
0 |
1 |
| Manual QA does not scale for AI. As traces grew to hundreds of megabytes with hundreds of tool calls, @retool turned to… |
@braintrust |
Company |
Original |
2026-03-25 |
449 |
2 |
0 |
1 |
0 |
2 |
| Use attachments in Braintrust to render SVG code as an image for human review evals. https://t.co/GBH2QGnXaC |
@braintrust |
Company |
Original |
2026-03-24 |
294 |
3 |
0 |
0 |
0 |
0 |
| https://t.co/5Zo6SD0QKX |
@braintrust |
Company |
Original |
2026-03-24 |
5,081 |
10 |
0 |
0 |
1 |
23 |
| A research paper found that using chain of thought reasoning improved LLM performance when detecting security vulnerabi… |
@braintrust |
Company |
Original |
2026-03-23 |
360 |
2 |
0 |
0 |
0 |
0 |
| Before Braintrust, @Replit relied on manual, multi-tool debugging to fix their AI agent's behavior.
Now @lhchavez and … |
@braintrust |
Company |
Original |
2026-03-20 |
729 |
8 |
0 |
2 |
0 |
3 |
| One dataset ≠ one eval.
It's possible to run multiple evals on the same dataset to test different measures of AI quali… |
@braintrust |
Company |
Original |
2026-03-19 |
293 |
4 |
1 |
0 |
0 |
0 |
| https://t.co/TgodJ8c7Bs |
@braintrust |
Company |
Original |
2026-03-18 |
56,082 |
67 |
4 |
1 |
1 |
313 |
| The best AI product managers know whether or not their product is improving.
DevRel engineer @darubberduckiee is hosti… |
@braintrust |
Company |
Original |
2026-03-17 |
406 |
6 |
1 |
3 |
0 |
0 |
| Starter is a new Braintrust pricing plan with no platform fee. It includes:
- 1GB processed data
- 10K scores
- 14 day… |
@braintrust |
Company |
Original |
2026-03-16 |
5,802 |
14 |
2 |
1 |
1 |
6 |
| What's new:
- Playground bulk optimization: annotate multiple outputs across prompts and let Loop optimize prompts for… |
@braintrust |
Company |
Original |
2026-03-13 |
892 |
6 |
2 |
1 |
0 |
0 |
| For providers subject to the GDPR, the architecture of a platform matters as much as the feature set it offers.
Braint… |
@braintrust |
Company |
Original |
2026-03-12 |
364 |
3 |
0 |
1 |
0 |
2 |
| Trace everything https://t.co/DJPXEX4TqD |
@braintrust |
Company |
Original |
2026-03-12 |
329 |
9 |
0 |
0 |
0 |
0 |
| https://t.co/5KJNYSeeMo |
@braintrust |
Company |
Original |
2026-03-11 |
13,023 |
24 |
1 |
0 |
2 |
70 |
| What's new in Braintrust: Observability and evals work best when built directly into developer workflows.
DevRel engin… |
@braintrust |
Company |
Original |
2026-03-10 |
439 |
6 |
1 |
2 |
0 |
0 |
| At @Dropbox, evals are the new PRD.
Evaluation scorers are defined before product development begins. Then production … |
@braintrust |
Company |
Original |
2026-03-09 |
604 |
8 |
0 |
1 |
0 |
2 |
| Trace fireside chat: Agent observability at scale.
@Replit built and scaled AI observability to support rapid iteratio… |
@braintrust |
Company |
Original |
2026-03-06 |
3,735 |
10 |
0 |
1 |
1 |
1 |
| Trace fireside chat: @Dropbox built Dash by treating evals like production code, using datasets, LLM judges, and automa… |
@braintrust |
Company |
Original |
2026-03-05 |
1,926 |
5 |
0 |
1 |
1 |
0 |
| Do you know what your AI is doing? https://t.co/gy1QKLilKu |
@braintrust |
Company |
Original |
2026-03-05 |
496 |
13 |
0 |
0 |
0 |
2 |
| The state of AI, in 2026: startup strategy pivots, why you can't vibe code a database, the sudden rise of sandboxes, an… |
@braintrust |
Company |
Original |
2026-03-04 |
684 |
7 |
3 |
5 |
0 |
0 |
| Notion’s evaluation practices have evolved from simple prompt-and-judge setups to a comprehensive framework that keeps … |
@braintrust |
Company |
Original |
2026-03-03 |
11,125 |
19 |
0 |
4 |
1 |
18 |
| Braintrust is hosting an afterparty at @humanxco with friends from @browserbase, @resolveai, @workos, @stainlessapi, @m… |
@braintrust |
Company |
Original |
2026-03-02 |
510 |
9 |
0 |
1 |
0 |
1 |
| Before running your first experiment, upload your dataset directly to Braintrust.
This unlocks cross-experiment compar… |
@braintrust |
Company |
Original |
2026-02-27 |
650 |
7 |
1 |
0 |
0 |
0 |
| That's a wrap on Trace.
Thank you to the speakers who shared their stories and the builders who joined us.
Now it's b… |
@braintrust |
Company |
Original |
2026-02-27 |
603 |
6 |
2 |
1 |
0 |
0 |
| Observability tells you what happened. Evals tell you whether it's getting better. Braintrust connects the two by autom… |
@braintrust |
Company |
Original |
2026-02-25 |
761 |
11 |
3 |
1 |
0 |
1 |
| Teams running AI applications in production generate thousands of traces a day, but can't read them all.
Braintrust is… |
@braintrust |
Company |
Original |
2026-02-25 |
866 |
9 |
3 |
1 |
0 |
2 |
| Trace is tomorrow. https://t.co/wZJNF4E33G |
@braintrust |
Company |
Original |
2026-02-24 |
4,097 |
14 |
1 |
1 |
2 |
1 |
| Looking for guidance on whether to start with Braintrust's free plan vs pro plan?
For simpler evals, the free plan goe… |
@braintrust |
Company |
Original |
2026-02-23 |
809 |
4 |
0 |
0 |
0 |
0 |
| To scale their voice agents across the globe, @navan built a continuous eval loop that observes production calls, under… |
@braintrust |
Company |
Original |
2026-02-23 |
836 |
9 |
0 |
1 |
0 |
2 |
| Trace is this week.
500+ AI builders. Hands-on workshops. Live demos. Sessions with the teams shipping quality AI. htt… |
@braintrust |
Company |
Original |
2026-02-22 |
1,125 |
11 |
0 |
0 |
2 |
3 |
| Custom views are useful when the feature you want isn't natively supported, because you can just implement it yourself.… |
@braintrust |
Company |
Original |
2026-02-20 |
705 |
6 |
1 |
0 |
0 |
2 |
| Trace-level scoring is here.
AI apps run as complex agents with tool calls and multi-turn conversations. Scoring one s… |
@braintrust |
Company |
Original |
2026-02-20 |
535 |
7 |
1 |
1 |
0 |
3 |
| Here's how to use the Braintrust MCP server to debug with minimal effort →
- Pull the trace after each eval run
- Chec… |
@braintrust |
Company |
Original |
2026-02-19 |
663 |
4 |
2 |
0 |
0 |
1 |
| Braintrust has raised an $80M Series B.
We're building the infrastructure that helps teams measure, evaluate, and impr… |
@braintrust |
Company |
Original |
2026-02-17 |
130,365 |
291 |
21 |
27 |
11 |
127 |
| Take human review on the go.
Human annotation and scoring is now optimized for mobile in Braintrust. https://t.co/MdE… |
@braintrust |
Company |
Original |
2026-02-14 |
873 |
5 |
1 |
0 |
0 |
1 |
| Coursera built a four-step approach to evaluating their AI features with Braintrust:
1. Define clear evaluation criter… |
@braintrust |
Company |
Original |
2026-02-13 |
527 |
3 |
0 |
1 |
0 |
0 |
| Opus 4.6 and GPT-5.3 just dropped and everyone wants to know which is better.
But don't obsess over the leaderboard ra… |
@braintrust |
Company |
Original |
2026-02-12 |
472 |
3 |
0 |
1 |
0 |
1 |
| How do you create an eval to judge how "good" AI-generated music is?
LLMs don't have ears, and music is subjective, so… |
@braintrust |
Company |
Original |
2026-02-11 |
385 |
4 |
0 |
1 |
0 |
0 |
| What are remote evals and why are they useful?
1. Some evals require complex dependencies (game engines, databases, sp… |
@braintrust |
Company |
Original |
2026-02-10 |
482 |
4 |
0 |
1 |
0 |
0 |
| What's new:
- Spans as table rows: filter experiments at the span level
- Autocomplete + linting: validate your code wh… |
@braintrust |
Company |
Original |
2026-02-09 |
472 |
5 |
1 |
1 |
0 |
1 |
| Evaluate how good Twelve Labs' language model Pegasus is at actually understanding what's going on in a video using Hug… |
@braintrust |
Company |
Original |
2026-02-06 |
825 |
5 |
2 |
1 |
0 |
1 |
| Trace talk → @modal
@bernhardsson will discuss the shift from shipping AI features to running AI systems, and what’s n… |
@braintrust |
Company |
Original |
2026-02-02 |
3,382 |
18 |
3 |
2 |
0 |
5 |
| Is bash really all your agent needs? We worked with @vercel to run head-to-head evals.
Result: tool choice matters, bu… |
@braintrust |
Company |
Original |
2026-01-22 |
343 |
7 |
0 |
0 |
0 |
0 |
| Trace talk → @replit
On Feb 25, @pirroh takes the stage to discuss agent observability at scale with @ankrgyl. If you'… |
@braintrust |
Company |
Original |
2026-01-22 |
1,308 |
11 |
4 |
0 |
0 |
1 |
| Build durable AI agents with @temporalio, get built-in observability with Braintrust. https://t.co/AwIaazd7cT |
@braintrust |
Company |
Original |
2026-01-20 |
739 |
3 |
0 |
1 |
1 |
1 |
| What's new:
- Thread layout search: search for keywords inside long LLM traces
- Raw trace search: search & downloa… |
@braintrust |
Company |
Original |
2026-01-16 |
3,698 |
7 |
0 |
1 |
1 |
3 |
| Ralph doesn't just eat paste, he eats tokens.
If you're going to let Ralph Wiggum run overnight, use Braintrust to log… |
@braintrust |
Company |
Original |
2026-01-15 |
366 |
3 |
0 |
1 |
0 |
0 |
| We've eliminated user-based pricing.
Invite more engineers, PMs, and cross-functional partners and build quality AI t… |
@braintrust |
Company |
Original |
2026-01-08 |
5,299 |
4 |
1 |
1 |
1 |
0 |
| Trace lineup is live.
Speakers from Ramp, Replit, Notion, Zendesk, Dropbox, HubSpot, FanDuel, Box, and more.
Bring yo… |
@braintrust |
Company |
Original |
2026-01-07 |
6,479 |
16 |
2 |
0 |
4 |
2 |
| Claude Code is fast for building agents. Debugging should be too.
We connected Claude Code and Braintrust as a two-way… |
@braintrust |
Company |
Original |
2025-12-23 |
1,028 |
4 |
0 |
1 |
1 |
3 |
| Brainstore is the purpose-built database for AI observability.
It handles the volume and complexity of modern AI syst… |
@braintrust |
Company |
Original |
2025-12-18 |
20,094 |
25 |
3 |
2 |
5 |
11 |
| What's new in Braintrust:
- Custom, vibe-coded trace views
- MCP servers
- Slack integration |
@braintrust |
Company |
Original |
2025-12-15 |
512 |
6 |
0 |
1 |
0 |
0 |
| Trace, our first user conference, is on Feb 25 at the California Academy of Sciences in SF.
500 AI builders. Hands-on … |
@braintrust |
Company |
Original |
2025-12-11 |
1,271 |
11 |
5 |
1 |
0 |
0 |
| You can now vibe code custom annotation UIs directly in Braintrust.
Investigate specific issues, share views across yo… |
@braintrust |
Company |
Original |
2025-12-08 |
510 |
3 |
0 |
1 |
0 |
0 |
| Production AI breaks in unexpected ways.
Loop helps you figure out what happened and what to fix, fast. https://t.co/… |
@braintrust |
Company |
Original |
2025-11-24 |
16,757 |
29 |
7 |
2 |
6 |
16 |
| Build AI that works. https://t.co/MUWJPbXhg2 |
@braintrust |
Company |
Original |
2025-11-21 |
1,471 |
7 |
2 |
0 |
1 |
3 |
| Gemini 3 is now available in Braintrust. https://t.co/NcWIhyZnb1 |
@braintrust |
Company |
Original |
2025-11-18 |
398 |
3 |
0 |
0 |
0 |
0 |
| See you at AWS re:Invent next month? We're hosting an afterparty on Wed, Dec 3 with @browserbase, @modal, & @llama_… |
@braintrust |
Company |
Original |
2025-11-10 |
1,161 |
5 |
2 |
1 |
0 |
0 |
| Update: @braintrustdata → @braintrust |
@braintrust |
Company |
Original |
2025-11-08 |
19,736 |
17 |
2 |
0 |
1 |
0 |
| From rebooking travel to managing expenses, @navan's AI voice agent tackles real-world challenges for travelers and fin… |
@braintrust |
Company |
Original |
2025-11-05 |
773 |
4 |
1 |
1 |
0 |
2 |
| There's nothing scarier than shipping without evals.
Learn how to evaluate even the spookiest images with Braintrust … |
@braintrust |
Company |
Original |
2025-10-31 |
3,692 |
7 |
1 |
0 |
1 |
4 |
| Another great Evals on Tap last night post @github Universe. Thanks for everyone that came by — see you at the next one… |
@braintrust |
Company |
Original |
2025-10-29 |
931 |
8 |
2 |
0 |
0 |
1 |
| Here's what's new in Braintrust:
- Spans in logs table
- "Pretty" data view
- Stream @Vercel traces to Braintrust
- Ja… |
@braintrust |
Company |
Original |
2025-10-25 |
724 |
4 |
0 |
1 |
0 |
0 |
| New on @readtechnically: Why evals are critical for AI that actually works.
Shows how teams evolve from "vibes-based" … |
@braintrust |
Company |
Original |
2025-10-24 |
595 |
2 |
0 |
1 |
0 |
0 |
| Human review is still a key part of many teams' eval workflows, especially when it comes to subjective definitions of q… |
@braintrust |
Company |
Original |
2025-10-23 |
568 |
2 |
0 |
0 |
0 |
1 |
| Non-technical teammates at @tolanworld spend hours a day in Braintrust making sure that everyone's favorite alien BFF f… |
@braintrust |
Company |
Original |
2025-10-22 |
6,919 |
14 |
2 |
1 |
2 |
8 |
| Join us next week for AI evals on Tap, an official @githubuniverse 2025 side event - featuring lightning talks on how r… |
@braintrust |
Company |
Original |
2025-10-21 |
475 |
3 |
1 |
0 |
0 |
0 |
| Braintrust is now available on the @Vercel Marketplace.
Run evals, monitor quality and user experience, and benchmark… |
@braintrust |
Company |
Original |
2025-10-16 |
515 |
8 |
0 |
1 |
0 |
0 |
| We're hosting the official @vercel Ship AI afterparty with our friends at Browserbase. Come hang out with other builder… |
@braintrust |
Company |
Original |
2025-10-15 |
6,936 |
14 |
2 |
1 |
3 |
0 |
| To run an eval, all you need is a task, dataset, and scorers.
When you put those together, you can expose regressions… |
@braintrust |
Company |
Original |
2025-10-13 |
447 |
2 |
0 |
1 |
0 |
1 |
| Here's what's new in Braintrust:
- Review queue for human feedback
- New custom monitor chart types
- JSON attachments… |
@braintrust |
Company |
Original |
2025-10-11 |
365 |
2 |
0 |
1 |
0 |
0 |
| Last night we brought together DX Engineers and Developer Marketers to discuss what makes a beloved developer brand. Th… |
@braintrust |
Company |
Original |
2025-10-10 |
855 |
6 |
2 |
0 |
0 |
0 |
| Our first #SFTechWeek event this week is Dev Tools on Draft with our friends at @graphite.
Thank you to everyone who … |
@braintrust |
Company |
Original |
2025-10-09 |
1,358 |
8 |
4 |
0 |
0 |
1 |
| Get started quickly with evals on Braintrust.
Learn how to build datasets, prompts, and scoring functions in less tha… |
@braintrust |
Company |
Original |
2025-10-07 |
494 |
5 |
0 |
0 |
0 |
2 |
| Last night @ankrgyl sat down with @turbopuffer, @cursor_ai, & @notionhq to discuss the real challenges behind AI en… |
@braintrust |
Company |
Original |
2025-10-03 |
1,034 |
11 |
0 |
0 |
0 |
0 |
| Next week is #SFTechWeek!
We're hosting two events for our developer and marketing friends in SF 👇️ |
@braintrust |
Company |
Original |
2025-10-02 |
763 |
3 |
1 |
1 |
0 |
0 |
| Have an AI application you want to add logging to but don't know where to start?
Our engineer Alex walks through how … |
@braintrust |
Company |
Original |
2025-09-23 |
462 |
4 |
0 |
0 |
0 |
1 |
| You can now retroactively apply scorers to historical data to find issues you missed.
Manually select from your logs, … |
@braintrust |
Company |
Original |
2025-09-22 |
466 |
6 |
0 |
0 |
0 |
0 |
| You can now analyze your experiments and logs with natural language.
Ask "show me errors from the last 24 hours" or "… |
@braintrust |
Company |
Original |
2025-09-17 |
417 |
3 |
0 |
1 |
0 |
0 |
| And we're just getting started. |
@braintrust |
Company |
Quote |
2025-09-10 |
540 |
8 |
0 |
0 |
0 |
0 |
| How do you build reliable AI tools at scale?
@withgraphite 's solution: systemic evaluation.
See how Graphite said go… |
@braintrust |
Company |
Original |
2025-08-26 |
4,849 |
11 |
2 |
1 |
2 |
3 |
| Async programming is the shift from line-by-line coding to problem-solving with AI agents.
By clearly defining proble… |
@braintrust |
Company |
Original |
2025-08-23 |
634 |
3 |
0 |
1 |
0 |
1 |
| We're grateful to be included in Forbes' Next Billion-Dollar Startups 2025 alongside many of our wonderful customers an… |
@braintrust |
Company |
Quote |
2025-08-12 |
2,636 |
16 |
2 |
1 |
0 |
1 |
| Braintrust is now available on AWS marketplace!
In addition to LLM evals and observability, you can access models from… |
@braintrust |
Company |
Original |
2025-08-11 |
821 |
5 |
1 |
1 |
0 |
0 |
| Hello from @Ai4Conferences 👋 |
@braintrust |
Company |
Quote |
2025-08-11 |
816 |
5 |
0 |
0 |
0 |
0 |
| Agents are transforming how we interact with technology, but building them can feel like navigating a maze of framework… |
@braintrust |
Company |
Original |
2025-08-08 |
57,690 |
46 |
2 |
1 |
2 |
65 |
| We'll be at @Ai4Conferences August 11-13 in Las Vegas - will you be in town?
Come say hi to team Braintrust at booth … |
@braintrust |
Company |
Original |
2025-07-29 |
1,337 |
4 |
1 |
0 |
1 |
0 |
| We've learned critical lessons from helping teams ship reliable LLM-powered products. Organizations using Braintrust ru… |
@braintrust |
Company |
Original |
2025-07-18 |
667 |
4 |
1 |
0 |
0 |
1 |
| In the world of AI development, the conversation often turns to eval frameworks, but we believe the real game changer i… |
@braintrust |
Company |
Original |
2025-07-16 |
577 |
3 |
0 |
1 |
0 |
0 |
| This Wednesday, 7/16, join us for a live session on how to evaluate agents.
We'll share how we built Loop: the AI age… |
@braintrust |
Company |
Original |
2025-07-14 |
871 |
5 |
0 |
0 |
1 |
2 |
| We put @xai's Grok 4 to the ultimate test: Can it create a good 'pelican riding a bicycle' SVG? (h/t @simonw)
Learn mo… |
@braintrust |
Company |
Original |
2025-07-12 |
1,092 |
5 |
0 |
0 |
1 |
1 |
| We're hiring! |
@braintrust |
Company |
Quote |
2025-07-12 |
1,546 |
12 |
0 |
1 |
0 |
1 |
| View eval traces in a friendly, chat-like UI with the new thread layout. https://t.co/Fe7gdOcXB7 |
@braintrust |
Company |
Original |
2025-06-12 |
975 |
7 |
1 |
1 |
0 |
1 |
| Evals should be easy.
Meet Loop, the AI agent for automatic prompt, dataset, and scorer optimization.
@aiDotEnginee… |
@braintrust |
Company |
Original |
2025-06-06 |
38,404 |
213 |
15 |
9 |
4 |
153 |
| evals |
@braintrust |
Company |
Original |
2025-06-05 |
672 |
5 |
0 |
0 |
0 |
0 |
| It's a great day to run evals.
@aiDotEngineer https://t.co/I4Z0UbdRxF |
@braintrust |
Company |
Original |
2025-06-04 |
4,324 |
20 |
3 |
7 |
0 |
3 |
| Come say hi to team Braintrust this week @aiDotEngineer world's fair! We're opposite Golden Gate Ballroom A.
We're al… |
@braintrust |
Company |
Original |
2025-06-03 |
3,582 |
10 |
1 |
1 |
1 |
0 |
| Join us at the AI engineer world's fair this week!
-Tuesday: 2 workshops on eval best practices w/ special guest @sara… |
@braintrust |
Company |
Quote |
2025-06-02 |
3,334 |
7 |
1 |
1 |
0 |
3 |
| Claude 4 is now available in Braintrust. |
@braintrust |
Company |
Quote |
2025-05-22 |
748 |
4 |
0 |
0 |
1 |
2 |
| 🧵 We shipped a UI refresh to make the Braintrust dashboard more intuitive and organized. |
@braintrust |
Company |
Original |
2025-05-20 |
767 |
5 |
0 |
1 |
0 |
0 |
| .@coursera uses structured AI evaluation to ship smarter features, like AI grading and a 24/7 learning coach, with conf… |
@braintrust |
Company |
Original |
2025-05-12 |
607 |
6 |
0 |
0 |
0 |
2 |
| Excited to support AI builders- don't forget to run evals before you ship the next big thing. |
@braintrust |
Company |
Quote |
2025-05-07 |
1,180 |
4 |
0 |
1 |
0 |
0 |
| You can now run remote Evals defined locally or hosted remotely directly in playgrounds.
Append --dev to npx braintru… |
@braintrust |
Company |
Original |
2025-05-07 |
761 |
2 |
0 |
0 |
0 |
0 |
| We're going on the road to NYC! Join us for AI Evals on Tap at NY Tech Week 2025.
Register here: https://t.co/7i9kW… |
@braintrust |
Company |
Original |
2025-04-29 |
827 |
4 |
1 |
0 |
0 |
1 |
| You can now build chains of prompts in the playground and run them consecutively.
Check out the docs for more info: … |
@braintrust |
Company |
Original |
2025-04-18 |
5,215 |
12 |
1 |
1 |
1 |
6 |
| We've added MCP server integration to help you debug and improve your app through natural language.
Add a few lines t… |
@braintrust |
Company |
Original |
2025-03-27 |
853 |
4 |
0 |
0 |
0 |
2 |
| You can now build with @OpenAI's o1-pro without writing any Responses API-specific code.
When you use our AI proxy, we… |
@braintrust |
Company |
Original |
2025-03-26 |
934 |
2 |
0 |
0 |
0 |
2 |
| Excited to share that we've been named to the Enterprise Tech 30 list for the second year in a row!
We're honored to … |
@braintrust |
Company |
Original |
2025-03-25 |
3,917 |
20 |
2 |
1 |
1 |
1 |
| Trace all your @OpenAI Agents SDK calls with one line of code! https://t.co/mNxVANtbpY |
@braintrust |
Company |
Original |
2025-03-11 |
18,630 |
103 |
13 |
5 |
2 |
46 |
| Braintrust is now 80x faster than any other LLM observability platform on the market.
To achieve this benchmark, we bu… |
@braintrust |
Company |
Original |
2025-03-03 |
99,416 |
200 |
13 |
8 |
15 |
109 |
| How does GPT-4.5 compare to your current model in production?
Evaluate both models side-by-side on Braintrust to find… |
@braintrust |
Company |
Original |
2025-02-27 |
9,604 |
2 |
0 |
0 |
0 |
1 |
| Grab your dataset from your production logs, pull it into a playground, and test @AnthropicAI's Claude 3.7 Sonnet side-… |
@braintrust |
Company |
Original |
2025-02-24 |
1,949 |
26 |
6 |
1 |
0 |
2 |
| New cookbook: Evaluating video QA
LLMs are great at interpreting text, but understanding and providing reliable answer… |
@braintrust |
Company |
Original |
2025-02-19 |
2,532 |
21 |
2 |
1 |
0 |
0 |
| If you use the Braintrust AI proxy, you'll be able to switch your production model to @grok 3 as soon as it's available… |
@braintrust |
Company |
Original |
2025-02-18 |
29,043 |
1 |
0 |
0 |
0 |
0 |
| We now fully support Amazon Bedrock and Google Vertex AI models via the playground and AI proxy.
Plus, you can now use… |
@braintrust |
Company |
Original |
2025-02-17 |
468 |
4 |
0 |
0 |
0 |
1 |
| But if you still want cookbooks, they're here for you.
https://t.co/KWGHNeKezi https://t.co/FEB5ifX7bV |
@braintrust |
Company |
Original |
2025-02-14 |
1,244 |
2 |
1 |
0 |
0 |
1 |
| https://t.co/uPXaV2kltI |
@braintrust |
Company |
Original |
2025-02-14 |
477 |
3 |
0 |
0 |
0 |
0 |
| New cookbook: Evaluating a voice agent
Learn how to simulate multilingual customer support calls, classify them, and m… |
@braintrust |
Company |
Original |
2025-02-13 |
444 |
5 |
0 |
0 |
0 |
2 |
| Building a multi-turn chat assistant?
Don’t forget to run evals:
https://t.co/iYspPNWQYF |
@braintrust |
Company |
Original |
2025-02-11 |
492 |
7 |
0 |
0 |
1 |
0 |
| Braintrust now supports structured outputs for most LLMs, including Anthropic and Gemini models.
Try it out with our n… |
@braintrust |
Company |
Original |
2025-02-10 |
355 |
2 |
0 |
0 |
0 |
0 |
| We had a blast at @spc talking about AI agents:
- How to evaluate them
- Best practices when building them
- Sharing in… |
@braintrust |
Company |
Original |
2025-02-07 |
403 |
3 |
0 |
0 |
0 |
1 |
| DeepSeek is the ‘LLaMa moment’ for O1-style (reasoning) models.
Read more from @ankrgyl and other AI infra leaders be… |
@braintrust |
Company |
Quote |
2025-02-06 |
723 |
3 |
0 |
0 |
0 |
0 |
| o3-mini is now available via the Braintrust AI proxy. |
@braintrust |
Company |
Quote |
2025-01-31 |
707 |
6 |
0 |
0 |
0 |
2 |
| New cookbook: Evaluating a prompt chaining agent
To produce production-ready agents, you need to understand what's goi… |
@braintrust |
Company |
Original |
2025-01-31 |
1,614 |
3 |
1 |
0 |
1 |
7 |
| LIVE NOW: How to evaluate AI agents
Join here: https://t.co/DbMveiqBcU |
@braintrust |
Company |
Original |
2025-01-30 |
744 |
3 |
1 |
1 |
0 |
0 |
| Ready for an inside look into building Zapier Agents?
Join us tomorrow, January 30 at 9AM PT to learn from Braintrust… |
@braintrust |
Company |
Original |
2025-01-29 |
650 |
5 |
1 |
2 |
0 |
0 |
| New blog post: Evaluating agents
Last month, @AnthropicAI published "Building effective agents," where they defined c… |
@braintrust |
Company |
Original |
2025-01-27 |
1,053 |
5 |
0 |
0 |
1 |
5 |
| Wondering how leading companies like @zapier are actually building and evaluating AI agents in production?
Reserve yo… |
@braintrust |
Company |
Quote |
2025-01-23 |
616 |
3 |
0 |
0 |
0 |
0 |
| Evaluate R1 against your own use cases on Braintrust. |
@braintrust |
Company |
Quote |
2025-01-21 |
673 |
4 |
0 |
0 |
0 |
0 |
| New cookbook: Building reliable multi-label classifiers
LLMs love making up labels—a huge problem when you need them t… |
@braintrust |
Company |
Original |
2025-01-17 |
479 |
3 |
0 |
0 |
0 |
2 |
| Building good AI agents is hard!
On Thursday, January 30th @ 9AM PST, we’re hosting a live webinar with Braintrust CEO… |
@braintrust |
Company |
Original |
2025-01-16 |
1,603 |
2 |
0 |
0 |
2 |
2 |
| 2024 was a crazy year. We overshot every goal we had thanks to an amazing community of customers who enable and inspire… |
@braintrust |
Company |
Original |
2025-01-01 |
1,116 |
10 |
0 |
1 |
0 |
6 |