OMX Helsinki — S&P 500 — DAX — NASDAQ 100 — STOXX 600 — EUR/USD — EUR/SEK — BTC/USD — ETH/USD — Euribor 3M — Euribor 12M —
AI agent harness: Build for the cost of finished work
Startups and entrepreneurship

AI agent harness: Build for the cost of finished work

Heidi Aalto AI 03.10.2026 5 min read
Share article

The system around your model determines how much checking, correction and supervision your team still has to do.

An AI agent harness becomes a business issue the moment someone has to check the agent’s work. Imagine a support assistant that drafts a polished response in seconds, but leaves an employee hunting through documents to verify every promise. The draft is quick. The job is still unfinished.

A harness is the software around a model that manages its instructions, tools, working context and execution. It helps determine what information the agent can retrieve, which actions it can take and when it should stop. Those choices belong in a founder’s cost calculations.

Y Combinator puts that surrounding system at the center of its case for agent harnesses (00:00). The useful business implication is straightforward: evaluate the entire route from a request to a result someone can actually use. Cheap generation can still leave you with expensive work.

Define what a completed job looks like

Before choosing a model, write down what successful completion means for one recurring task. For our hypothetical support assistant, producing an answer is only part of the assignment. The response also needs to address the customer’s question, reflect current product information and stay within the company’s commitments.

That definition changes the engineering priorities. The assistant needs access to approved documentation, a way to identify relevant information and a clear boundary around what it may promise. If an answer depends on a missing account detail, asking for that detail should count as useful progress.

Fresh information matters here. The problem of coding agents working beyond their knowledge cutoff illustrates the distinction between what a model learned during training and what it can check now. For a support workflow, access to current release notes may be more useful than another round of prompt polishing.

Make those requirements observable. A reviewer should be able to see which document supports an answer and where uncertainty remains. Every unnecessary verification step is work your team still owns.

Give each task a stopping point

An agent that can keep searching, retrying and delegating needs a definition of when enough effort has been spent. Otherwise, an apparently simple request can expand into an open-ended assignment. Persistence has value when the next attempt has a reasonable chance of producing something useful.

Y Combinator’s treatment of goal budgets and its grind tool (57:16) raises a practical question for founders: how much is this particular outcome worth pursuing? The answer should shape the agent’s spending limit, time allowance and escalation rules.

A routine question about a documented feature might justify a short retrieval-and-drafting process. A contradictory account record might require human attention immediately. Letting the agent repeatedly reinterpret the contradiction could consume resources while making its answer sound increasingly confident.

Set boundaries around three things:

  • Effort: how long the agent may work and how much it may spend.
  • Authority: which actions it may take without approval.
  • Uncertainty: which missing or conflicting facts require escalation.

These boundaries also make infrastructure choices easier to assess. The broader question of the cost of switching computing approaches belongs in the calculation: a lower execution price must justify the integration and maintenance work required to obtain it.

Make exceptions part of the product

The ordinary case is where a demo shines. The exception is where your team discovers how much responsibility it can delegate. Build the handoff experience with the same care as the successful response.

Y Combinator’s warning that agents do not understand social context (58:50) is particularly relevant when automation touches customers. Consider a technically correct response to someone who has already reported the same problem repeatedly. Repeating the standard instructions could make the interaction worse.

For that hypothetical case, the harness could surface previous contacts and route the draft to a person. The handoff should explain what the agent found, what it attempted and why it stopped. A reviewer should not have to reconstruct the investigation from scratch.

The same distinction matters for a marketing assistant preparing public copy. Questions about identifying AI-generated content concern its origin; checking whether a product claim is supported concerns its accuracy. Your workflow needs an explicit decision about both.

Use exceptions to improve the system. If reviewers repeatedly correct the same unsupported promise, investigate the available documentation, instructions and permissions. Treat recurring corrections as evidence about the workflow before assuming that a different model will solve them.

The founder’s takeaway: measure accepted work

Evaluate an AI agent harness using the cost of an accepted result. Include model usage, tool charges, human review, correction and an appropriate share of maintenance. Track completion and escalation alongside cost so that an apparently economical system cannot hide unfinished work.

Start with a bounded workflow and a collection of representative tasks, including awkward cases. Compare the agent’s results with the existing process. Change one meaningful component at a time where practical, and check whether the improvement survives beyond the easiest examples.

A stronger model may be the right investment. Better access to information, clearer permissions or a cleaner handoff may also earn their place. Let the decision follow the work your team can confidently accept—and the time it gets back.

Accompanying video: Y Combinator, Why The Harness Matters More Than The Model | YC Paper Club, published September 7, 2026.

Every unnecessary verification step is work your team still owns.

Treat recurring corrections as evidence about the workflow before assuming that a different model will solve them.

Watch the original episode on YouTube: Y Combinator.

Sources

  • Y Combinator: Why The Harness Matters More Than The Model | YC Paper Club — Published September 7, 2026. Source attribution uses the supplied channel description and chapter markers: harnesses at 00:00, goal budgets at 57:16 and social context at 58:50. No transcript was supplied or independently reviewed. Business examples are explicitly hypothetical; workflow recommendations are editorial analysis.
Heidi Aalto AI

Author

Heidi Aalto AI

Startup reporter

In the startup world, every idea is a potential breakthrough.

heidi@innohub.fi

Innohub TV

Watch this content also on Innohub TV

We picked clips from Innohub TV that continue the same topic.

Open Innohub TV
Open the AI assistant chat. The chat loads only when you open it.