Skip to content

The two-line estimate: generation, verification, and which one is bigger

Your buyer already assumes AI made the build cheaper, and an hourly quote invites them to cut it line by line. A quote with two lines, generation and verification, shows where the saved hours went before anyone asks. In our experience they went into review, which is now the larger line.

pricing · procurement · estimates

Your buyer has already decided that AI made your work cheaper. The question is no longer whether it did. It is where the savings went and why they are not in the price. Our answer is a quote with two lines instead of a column of hours: generation, which did get cheaper, and verification, which did not. In our experience, verification is now the bigger line.

The question is already on the table

On the procurement side, the question is put directly: where did AI make delivery cheaper for the supplier, and how does that show up in the price? According to Digital Applied's pricing guide (June 2026), the CPOs who are winning take it into every renewal.

Agencies are hearing it too. The same guide cites a 2025 Productive.io survey of 180+ agencies: roughly a third had already received explicit requests for an "AI discount", and about half expected them soon. At the top of the market, WPP ties 20–25% of its net sales to performance-linked fees.

One caveat. Every one of these numbers comes from someone who sells a pricing model or a service, and the 2025 survey reaches us second-hand. We read them as a signal of what buyers are asking, not as a measurement of the market.

Where the saved hours actually went

The buyer's assumption is half right. Generation is faster. What it misses is where the work moved.

Faros AI's 2026 telemetry report covers 22,000 developers across more than 4,000 teams and measures what changed after AI tool adoption. Quoting the report: "Median time in review is up 441.5%" and "Bugs per developer are up 54%". Note the unit: that second figure counts bugs per developer, not per pull request. Faros sells engineering analytics, so this is a vendor's dataset, and we cite it as one.

Writing code got faster. Checking it got slower, and there is more to find when you do. The work did not disappear. It changed places.

Our view: this is the line estimates break on. The build line holds up. The review line doesn't. An honest quote prices verification, because on the work we see, verification is now most of the work.

Why an hourly quote loses this argument

Our view: hourly billing puts the supplier's cost structure on the invoice, and every line on it becomes something for the buyer to negotiate down. Here is what the buyer sees:

Quote v1 — hourly
-----------------------------------------
Backend development        ...  h
Frontend development       ...  h
Integrations               ...  h
QA                         ...  h
Project management         ...  h
-----------------------------------------
Buyer's read: "AI writes code now. Cut every line."

Every line is a target, and none of them says what AI changed. In our experience, review time in a quote like this hides inside QA, and QA looks like the easiest line to trim. In our reading of the telemetry above, review is exactly the work that grew.

The two-line version

Here is the same work, split by what AI changed and what it did not:

Quote v2 — two lines
-----------------------------------------
1. Generation
   Scaffolding, first drafts of code, tests, migrations.
   AI made this cheaper. The saving is shown here.

2. Verification
   Review of every change by someone who did not write it.
   Validation re-run after each fix.
   Defect triage, regression checks, acceptance against the brief.
   AI made this larger. This is where the saved hours went.
-----------------------------------------
Buyer's read: "I can see the discount, and I can see what it bought."

The first line answers the procurement question before anyone asks it. The second explains why the total did not fall by the same amount. A buyer who arrives expecting a discount sees one, on the line where it is real.

This part is our opinion, not a finding: a quote in this shape is harder to negotiate line by line, because the lines are no longer interchangeable hours. Cutting line 2 no longer means "trim QA". It means "accept less review while bugs per developer are rising", and the telemetry above shows what that looks like.

What goes in the verification line

"Verification" only survives a procurement review if it is specific. In our practice it breaks down like this:

Verification — what it contains
-----------------------------------------
review      a fresh reviewer, no shared context with the author
exit rule   validation green + zero open findings on the current state
evidence    tests re-run in full after each fix batch, never shrunk
scope       every changed file covered; an empty scope is refused
record      each finding numbered, fixed on its own branch, merged after review

We run this as a loop, not a checkpoint. Every finding gets a number, is fixed on its own branch and is merged only after review. That numbered record is the verification line made visible: a buyer can see it itemized instead of taking a QA estimate on trust.

Where AI made our work cheaper

In our experience, first drafts of code, boilerplate, test scaffolding and migrations are faster to produce than they used to be. That saving is real, and it belongs on line 1.

What did not get cheaper: understanding what the code does, checking it against the brief, and owning it when it breaks. We would rather name that trade-off than bury it in an hourly rate.

We do not publish a ratio between the two lines. It depends on the product, the risk and how much of the system already exists. On the work we take, line 2 is the bigger one.

FAQ

Doesn't a separate verification line just look like a way to keep the price up? It would if it were a single number. That is why it is itemized: who reviews, what the exit rule is, what evidence is re-run. A buyer can challenge any item. In our experience, a line like "QA, 40 h" gives a buyer nothing to challenge except the number.

How reliable is the 441.5% figure? It comes from Faros AI's 2026 telemetry report on 22,000 developers across more than 4,000 teams, which states: "Median time in review is up 441.5%". The 54% in the same report is "Bugs per developer", not defects per pull request. Faros sells engineering analytics, so treat it as a vendor's dataset rather than an independent study. In our view the mechanism holds either way: generation is cheap and review is not.

Should a studio just offer the AI discount? Our view: offer it where it is real, on generation. A discount spread across every line cuts verification too, and verification is where defects get caught.

Does this only work for fixed-price quotes? No. The two lines work with hourly, fixed or outcome-based pricing. What matters is what the buyer can see, not how the total is calculated.

What would change our mind? Measured data showing that AI-assisted delivery cuts review time and bug rates at the same time. If that shows up, line 2 shrinks, and we will say so here.

A related case

In Holt, a language-learning product we built, verification is designed into the product. No model output reaches a learner unvalidated: an independent validator checks every generated card before it is shown. Read the Holt case.

If you have something worth building, we'd like to hear about it.