Braintrust is an observability and evaluation platform for AI agents. It captures what your agent did in production and scores whether the output was any good. Failing cases become tests, so the same mistake doesn't ship twice.
It's a genuinely strong product. Whether it's the right product depends on what else you need it to do.
Two things before the review. First, there are two companies called Braintrust, and only one of them is this one. Second, the pricing jump is bigger than most reviews mention. Both of those are covered below.
Which Braintrust Are You Looking For?
This trips up almost everyone searching, so let's clear it immediately.
braintrust.dev is the AI observability and evaluation company. Founded 2020, San Francisco, backed by ICONIQ and Andreessen Horowitz. That's the one this article is about.
usebraintrust.com is a freelance talent network. Completely unrelated company, completely different product. If you landed here looking for freelancer pay rates or contract work, that's the site you want.
The two share a name, a first page of Google, and nothing else. Worth knowing before you go hunting for community opinion: most Reddit threads ranking for "braintrust" are about the talent network, not the eval platform.
What Is Braintrust AI?
Braintrust calls itself "the active observability platform for agents." The word doing the work there is active. It isn't just a place where traces land. The pitch is that you observe, evaluate, and then change the thing you're observing, on a loop.
The company was founded in January 2020 by Ankur Goyal, who is also CEO. Goyal previously founded Impira, which Figma acquired. Braintrust is headquartered in San Francisco.
On 17 February 2026 the company announced an $80 million Series B led by ICONIQ, with Andreessen Horowitz, Greylock, Elad Gil, and basecase capital participating. Braintrust's own announcement doesn't state a valuation. Press coverage at the time put the round at an $800 million valuation, so treat that figure as reported rather than company-confirmed. Andreessen Horowitz published its own investment rationale if you want the investor framing.
The customer list is the strongest signal in the whole review. Notion, Stripe, Vercel, Instacart, Zapier, Replit, Ramp, Dropbox, Cloudflare, Coursera, Graphite, and Navan all use it. Notion's case study is the one worth reading: across 70 engineers, building evals into the workflow took them from fixing roughly 3 AI issues a day to 30. That's a tenfold change in shipping speed, from a company that ships a lot.
What Does Braintrust Do Well?
The docs are organised around six things: instrument, observe, annotate, evaluate, deploy, admin. In practice four of them are where the product earns its price.
- Tracing. Every step of the agent's reasoning gets captured: prompts, tool calls, retrieved context, latency, cost. Traces land in Brainstore, a database Braintrust built specifically for this workload. Querying millions of traces stays fast, which is not a given with trace data.
- Evaluation. This is the signature. Score outputs with built-in scorers, LLM-as-a-judge, custom code, or human review. If your agent answers from documents, the scorers worth writing first are the ones checking retrieval quality. Evals run in CI/CD, so regressions get caught before release. One click turns a production trace into a test case, and that single mechanic is the reason teams stick with it.
- Discover. Automated pattern surfacing across your traces, grouped into Topics, with Custom Facets if you want to cluster by something specific to your business. This is newer, and it's aimed at the "I have 40,000 conversations and no idea what's wrong with them" problem.
- Loop. An AI assistant that reads your traces and proposes better prompts, scorers, and datasets. Describe what you want to optimise in plain language and it generates it. If you've done this by hand, it's prompt chaining with the iteration automated.
Two smaller things that matter more than they look. There are task-specific trace views, so an annotation interface can be shaped like the workflow being reviewed rather than a generic tree. And there's MCP support, which means you can pull traces into your IDE and work on them where you already are. If your team lives in Cursor or Claude Code, that shortens the loop meaningfully. (If MCP is new to you, our prompt-engineering guide covers where it fits.)
What Does Braintrust Cost?
Three tiers, and the middle one is where people get surprised.
Starter is $0. You get $10 in credits, 1 GB of processed data, 10,000 scores, and 14-day retention. Unlimited users, projects, datasets, and experiments. It's enough to evaluate the product. It is not enough to run anything real, and calling it generous overstates it.
Pro is $249 per month, which buys $249 in credits, 5 GB processed data, 50,000 scores, 30-day retention, plus custom charts, environments, RBAC, and priority support.
Enterprise is custom, with negotiated retention, export, RBAC, and on-prem or hosted deployment.
Model two things before you commit. Processed data and scores are the meters, and agentic workloads move both faster than chat workloads do. A single multi-step agent run generates far more trace volume than one request-response exchange, because every reasoning step, tool call, and retrieval is its own unit of billable data. If you're modelling cost from token counts alone, you'll be low. And retention is priced separately from volume. Fourteen days on the free tier means an intermittent bug you notice on a Monday may have already aged out. Check the current pricing page before you budget, since tiers move.
Where Braintrust Falls Short
No tool is perfect. Here's what to actually watch for.
Braintrust doesn't build agents. It watches them. You build your agent somewhere else and send traces in. That means your building tool and your observability tool are two different systems with two different mental models. You spot the problem in one and fix it in the other. For teams with dedicated AI engineers, that's a reasonable split. For teams where the person who notices the problem is also the person who has to fix it, it's friction on every single iteration.
The Pro tier is a real jump. Free to $249 is a big step for a bootstrapped team. Open-source options like Langfuse or Arize Phoenix give you more at that stage, in exchange for running the thing yourself.
No runtime guardrails. Braintrust scores outputs after they happen. It has quality gates that block bad releases, which is not the same as blocking a bad response. If you need something intercepting unsafe output before a customer sees it, you'll add a second tool. Galileo does this natively, with real-time checks on inputs and outputs during live requests. Our guide to preventing LLM hallucinations covers the techniques that reduce the need for a net in the first place.
It's engineering-centric. Annotation and review genuinely help bring non-engineers in. But tracing, evals, datasets, and experiments assume someone comfortable with SDKs and CI/CD. A product manager can participate. An engineer still drives.
No conversation-level simulation before launch. This is the claim most reviews get wrong in both directions, so to be precise: Braintrust does have environments: dev, staging, and production, with version pinning plus webhook or Slack alerts when a version gets promoted. What it doesn't have is a way to simulate thousands of conversations across different personas before the agent goes live. Environments version your prompts and parameters. They don't rehearse the conversation.
The community is smaller, though "closed-source" is wrong. You'll read that Braintrust isn't open-source. That's not accurate. autoevals has around 990 stars, braintrust-proxy about 408, agentbehavior about 244, and the Python and JS SDKs are MIT and Apache licensed. The platform is closed. The eval tooling isn't. That said, the surrounding community is smaller than Phoenix's, which crossed 10,000 stars in June 2026, or Langfuse's or LangSmith's. Fewer third-party tutorials, fewer community integrations, a smaller pool of people who have hit your exact problem.
Who Shouldn't Use Braintrust?
Four cases where it's the wrong pick:
- You need one platform that designs, deploys, and observes. Braintrust covers the last part only.
- Your team is mostly product managers, CX leads, or conversation designers. You'll need engineering support to get value.
- You need real-time guardrails that block bad output before it reaches a user. Braintrust evaluates after the fact.
- You're building support or lead-gen agents and need to see the conversation inside the workflow that produced it, not as a trace tree.
That last one is worth expanding, because it's the difference between a debugging tool and a fixing tool.
How Braintrust Compares to the Alternatives
Braintrust is the strongest option on this list if what you need is eval rigour on an agent you're building yourself in code. That's a real category and it wins it. The table is here because "best" depends entirely on what else you need.
| Platform | Best for | Watch out for |
|---|---|---|
| Braintrust | Eval-first rigour, CI/CD gating, trace-to-test-case workflow | $0 → $249 jump; no runtime guardrails; separate from wherever you build |
| Langfuse | Open-source, self-hostable, cheap at small scale | You operate it; less opinionated about eval methodology |
| LangSmith | Teams already deep in LangChain or LangGraph | Pulls you toward one framework's abstractions |
| Arize Phoenix | Classical ML plus LLM under one roof; large OSS community | AX pricing meters on spans and burns fast on agents; thin pre-production |
| Galileo | Real-time guardrails and policy enforcement during live requests | Guardrail-led rather than eval-led; less of a dev-loop tool |
| Voiceflow | Building, deploying, and observing the same agent in one place | Not a general-purpose LLM eval framework for arbitrary code |
For a wider view of the category, see our breakdown of AI agent observability, the honest read on Arize AI, and the roundup of agent management platforms if governance is the part you're solving for.
How Voiceflow Approaches Observability Differently
Braintrust answers: was this output good?
Voiceflow answers that too, and then: why did the agent say it, where in the flow did it happen, and how do I fix it right now?
Voiceflow is an AI agent platform. You design, build, deploy, and observe in the same place. Observability isn't a system you pipe data into. It's how you iterate.
Transcripts show every conversation your agent had with a real user, turn by turn. Replay them. Inspect tool calls and model responses in context. When something goes wrong you're looking at the interaction the customer actually had, not a reconstruction of it.
Evaluations score those transcripts automatically against criteria you define: resolution rate, satisfaction, compliance, whatever matters to you. Pick the criteria carefully, since deflection rate on its own will flatter you. They run on every new conversation and can be applied retroactively to historical data.
Agent logs give engineers the span-level depth: tool calls, model decisions, knowledge base retrieval scores, timing per step, and the exact point where a conversation went to a human agent. Same detail you'd get from a trace tree, sitting next to the flow that produced it.
Analytics give product leads and execs a dashboard that answers "how are our agents doing" without an engineering ticket. That's usually the same audience asking about ROI.
Environments are dev, staging, and production for the whole agent, not just its prompts. That's the honest distinction from Braintrust's environments: same idea, different unit of versioning.
Voiceflow is model-agnostic, covering OpenAI, Anthropic, Google, or your own. So an eval that says a model is underperforming becomes a change you can actually make. And it's SOC 2 Type 2 with PII masking, which matters when the traces you're storing contain customer conversations. Our security and compliance guide covers what to ask any vendor here.
Teams running this in production include Turo, StubHub International, Sanlam Studios, and Trilogy, where AI now handles 59% of support conversations end to end. If you're moving from a pilot to something real, the 90-day path from pilot to production is the framework we use, and replacing a legacy chatbot is the version of that story most enterprises are actually living, and omnichannel support covers what changes when the same agent runs across channels.
So Is Braintrust the Best for AI Observability?
If "best" means best at evaluating LLM output quality, Braintrust has a real claim. The eval-first approach is correct. The production-trace-to-test-case mechanic is genuinely fast. The customer list speaks for itself.
But evaluation is half the job.
The other half is what happens after you find the problem, and that's where the separate-tool model costs you. You spot a regression in Braintrust. You switch to your framework. You find the code. You change it. You redeploy. You wait for new traces. Then you check whether it worked.
In Voiceflow you see the problem in a transcript, click through to the flow that produced it, adjust the logic, test in staging, promote to production, and watch the evaluation move. Same session.
The fastest path to a better agent isn't more dashboards. It's less distance between noticing and fixing.
Braintrust shortens the eval loop. If you also want the fix loop shortened, you're looking for a different shape of tool.