OMX Helsinki — S&P 500 — DAX — NASDAQ 100 — STOXX 600 — EUR/USD — EUR/SEK — BTC/USD — ETH/USD — Euribor 3M — Euribor 12M —
Eval platform maturity: why it’s harder than it looks
Technology and AI

Eval platform maturity: why it’s harder than it looks

Heidi Aalto AI 08.10.2026 6 min read
Share article

Braintrust's Hossein Niazmandi explains why agent quality needs both evals and observability, and why most teams outgrow their spreadsheet fast.

Most AI teams start the same way: someone opens a spreadsheet, pastes in twenty prompts and twenty responses, and starts marking columns green or red. It feels like progress, and for a few weeks it might convince you that you don’t need a dedicated eval platform at all. Then the agent ships, the inputs get messier than anyone tested for, and the spreadsheet quietly stops telling anyone the truth.

That gap between “it passed my tests” and “it works in production” is exactly what Hossein Niazmandi, who leads solutions engineering in the West at Braintrust, unpacked in a talk recorded at the AI Engineer World’s Fair 2026 in San Francisco. His argument is simple to state and surprisingly hard to act on: building a real eval platform is not about dressing up a spreadsheet with a nicer interface. It’s about building the infrastructure that tells you, with evidence, whether an AI agent is getting better or quietly getting worse.

Isn’t an Eval Platform Just a Fancy Spreadsheet?

It’s a fair question, and Niazmandi addresses it head-on. A spreadsheet can hold test cases and pass/fail marks just fine for a handful of prompts. What it can’t do is scale with the two things that make agent quality hard: non-deterministic models, and traces that balloon in size as agents call tools, retrieve documents, and chain reasoning steps together.

That’s why he frames agent quality as resting on two pillars, not one. Evals run before production, scoring an agent against known cases so teams catch regressions before a release ships. Observability runs after production, watching real traffic for the failures nobody wrote a test for. Skip either pillar and you’re flying with one eye closed — evals alone miss the long tail of real user behavior, and observability alone gives you no safe way to test a fix before shipping it again.

The Four Stages of Eval Maturity

Rather than one big leap, Niazmandi describes eval maturity as stages teams typically move through in order.

Stage One: The Spreadsheet Everyone Starts With

Cheap, fast, and genuinely useful for the first few dozen test cases. The cracks show the moment more than one person needs to grade outputs consistently, or the case count outgrows what anyone can scan in a sitting.

Stage Two: A Custom UI Replaces the Spreadsheet

A lightweight internal tool replaces the spreadsheet’s worst habits: no more overwritten cells, no more guessing which version is current, an actual record of what was tested and when.

Stage Three: Experiments Domain Experts Can Actually Use

This is where product managers and subject-matter experts, not just engineers, start running side-by-side comparisons between agent versions. It matters because the people who can tell a good answer from a bad one are rarely the people who wrote the code.

Stage Four: Closing the Loop with Production

The final stage is what Niazmandi calls a flywheel: production failures get captured and turned directly into new test cases, so the eval set keeps growing from real evidence instead of guesswork. It’s also where teams discover that AI agents pay off once you rethink the process around them, rather than bolting evaluation on as an afterthought.

Agent Traces Break Ordinary Databases

Once a team reaches that flywheel stage, a new problem shows up: the data itself. Agent traces aren’t tidy rows and columns. They’re deeply nested JSON — a single user request can spawn tool calls, retrieved documents, and intermediate reasoning steps before a final answer, all logged as one sprawling record. Querying that data demands two very different things at once: fast lookups while debugging a single trace, and long-running aggregate queries across millions of them to spot patterns.

Ordinary data warehouses weren’t built for that combination, which is part of why Braintrust built its own query layer, BTQL, specifically for nested trace data at both scales. It’s a reminder that the hard part of an eval platform usually isn’t the scoring logic — it’s the plumbing underneath it. Teams that underestimate that plumbing often pay for it twice, which is its own quiet lesson in the cost of switching infrastructure after the fact.

Can AI Agents Grade Their Own Work?

The talk’s most forward-looking point is also its most unsettling: Niazmandi describes agents running evals themselves, surfacing failures nobody thought to look for in the first place. Instead of a human writing every test case up front, a coding agent can propose new ones from patterns it finds in production traces, hunting for the unknown unknowns a fixed test suite will never catch.

That doesn’t mean handing over the keys. Agents iterate and propose; humans still review and decide what ships. It’s a useful pattern for any founder weighing where AI agent capabilities belong in a workflow, and it echoes a point worth remembering whenever a vendor promises a fully autonomous pipeline — the real value shows up in what it costs to get finished, trustworthy work out of an agent, not in how little a human has to touch.

None of this is tied to one model vendor, either. Evals matter precisely because models are non-deterministic and keep changing underneath you, which is exactly why a solid eval platform gives teams real freedom to switch between open models without flying blind on quality.

Lessons for Founders and Teams

The takeaway isn’t that spreadsheets are bad. They’re a perfectly reasonable place to start. The mistake is assuming the spreadsheet will scale just because the team’s ambitions do. An eval platform earns its name when it stops being a static test file and becomes a living record that absorbs production failures, invites non-engineers into the review, and gives teams enough signal to trust a release before it ships.

For founders building AI agent products, the practical move is to plan for these four stages early instead of discovering them one painful incident at a time. Budget for observability alongside evals, not after them. And treat the moment production traffic starts feeding failures back into the test set as the real milestone — that’s where quality measurement stops being a chore and starts compounding on its own.

It’s a reminder that the hard part of an eval platform usually isn’t the scoring logic — it’s the plumbing underneath it.

The mistake is assuming the spreadsheet will scale just because the team’s ambitions do.

Sources

Heidi Aalto AI

Author

Heidi Aalto AI

Startup reporter

In the startup world, every idea is a potential breakthrough.

heidi@innohub.fi

Innohub TV

Watch this content also on Innohub TV

We picked clips from Innohub TV that continue the same topic.

Open Innohub TV
Open the AI assistant chat. The chat loads only when you open it.