Back to articles

Fixed Scope vs Flat Rate Software Development: A 2026 Rubric

August 25, 2026
Fixed Scope vs Flat Rate Software Development: A 2026 Rubric

Fixed scope vs flat rate software development, decided on evidence. Why estimates lost their reference class in 2026, and a rubric for picking the contract.

Share:LinkedInX

The Estimate Stopped Being a Measurement

In July 2025, METR ran the most rigorous controlled study anyone has published on AI and developer productivity: 16 senior open-source maintainers, 246 real issues, in repositories they had owned for years. Tasks took 19% longer with AI, confidence interval +2% to +39%, while the developers themselves believed they had gone 20% faster.

The same lab re-ran it later that year with the developers who came back. The new estimate is that AI made them roughly 18% faster — confidence interval −38% to +9%, which crosses zero and is therefore not statistically significant. METR's own figure caption is careful about it: "Late-2025 AI likely accelerated open-source developers, but selection effects obscure the true speedup."

In February 2026 they announced they were scrapping the experimental design. Not because of the result — because they can no longer build a control group. Developers increasingly refuse to take the no-AI arm at all.

Pick whichever number you find more credible. The part worth noticing is the shape: a dedicated research lab, twelve months, a point estimate that moved about 37 points, and a method that stopped working underneath them.

That is not a story about whether AI helps. It is a story about reference classes — and it is the reason the fixed scope vs flat rate software development question stopped being a matter of taste and started being a matter of evidence.

Estimation Is Extrapolation, and the Class Stopped Holding Still

An estimate is an analogy. This looks like the thing we did last quarter; that took two weeks; call it two weeks.

Story points, t-shirt sizes and planning poker are the crude version of that move. Probabilistic forecasting — Monte Carlo over historical cycle times, quoted in percentiles — is the good version, and it is genuinely more robust, because it never needed per-item accuracy in the first place. The flow-metrics community has been right about this since roughly 2011.

But every method in that family rests on one assumption: that recent history is a usable sample of the near future. That the cycle-time distribution is roughly stationary. Nothing in the toolkit survives a base rate that moves 37 points in twelve months, in a direction the best-funded lab in the field cannot pin down, landing differently depending on which part of the codebase the ticket falls into.

What the 2026 Telemetry Shows

Faros AI's 2026 engineering report, published in April, covers 22,000 developers across 4,000+ teams and two years of pipeline telemetry. Worth naming the interest: Faros sells engineering-intelligence tooling, and a report concluding "you cannot see your own downstream decay" is convenient for them. The underlying data is still pipeline telemetry rather than a survey, which is more than most of what gets cited on this subject.

Output went up. Tasks per developer +33.7%. Epics per developer +66%. AI code acceptance climbed from 20% to 60%.

Now the downside, with denominators, because the headline versions of these numbers circulate without them:

  • Bugs per developer rose 54%. But tasks per developer rose 33.7% over the same period, so bugs per task rose about 15%. Measured against epics it is roughly flat. The 54% figure is real and it is misleading.
  • The number that survives normalisation is the incidents-to-PR ratio: +242.7%. Per unit of work merged, production incidents roughly tripled. That one is not a denominator artefact.
  • Median time in review: +441.5%. And 31.3% more pull requests merged with no review at all. Review capacity did not grow. Some of it was skipped.
  • Code churn under high adoption: +861%, alongside acceptance climbing from 20% to 60%. Those are not two findings. That is one loop — accept more, rewrite more.

On a subset Faros could measure, roughly 10% of the dataset and explicitly flagged as high-variance, lead time from commit to production rose around 480% while deployments per week fell about 11%. We would not build an argument on a high-variance tenth of a dataset, so treat it as a hint. It points the same way as everything above it.

More code merged. Not more software shipped.

Why This Is a Forecasting Problem, Not Only a Quality Problem

Three numbers in the same report are about predictability rather than speed: average time a task spends in progress +225.2%, tasks stalled seven days or more +26%, work restarts +13.8%. Those are dispersion metrics. Not slower — less predictable.

And the effect is not uniform, which is the part that breaks planning. DORA's 2026 return-on-investment work cites Stanford data putting gains at 35–40% on greenfield against 10% or less on complex legacy code. Same engineer, same nominal ticket size: an afternoon in the new service, most of a week in the ten-year-old billing module. The variable that now dominates delivery time is one the estimate has never encoded, because until recently it did not dominate anything.

May's systematic review — arXiv:2605.01160, a preprint covering 67 sources from January 2022 to April 2026 — calls this the Productivity-Reliability Paradox, and names three moderating variables: task abstraction level, codebase maturity, developer seniority. Those are the three that no estimate contains.

DORA also offers the most serious argument against everything that follows: a J-curve, where organisations get worse before they get better. If that is right, this is a transition rather than a new equilibrium, and the answer is to invest in engineering foundations and wait it out. That is plausible. It is also an argument against signing a twelve-month fixed-scope commitment at the bottom of a J-curve.

The Contract Is the Last Thing to Move

Methodology moved fast — spec-driven development, agent harnesses, feedback sensors, most of it sitting in Assess on Thoughtworks' April 2026 Radar with agent swarms under Caution. Commercial models have not moved in a decade. Most custom software is still bought exactly as it was in 2015: discovery phase, fixed scope, signed statement of work, hourly rate, change-request process for everything discovery could not have known.

That apparatus encodes two premises — writing code is expensive, and changing your mind is expensive. The first clearly no longer holds. The second is now mostly an artefact of the process rather than of the work: a one-day change routinely takes a scoping call, a revised estimate, an approval cycle and an amended SOW. A week of calendar and a day of billable time on both sides, to authorise a day of work.

One honest complication. Anthropic's 2026 agentic coding report finds that roughly 27% of AI-assisted work is work that would not have been done otherwise. We read that as scope becoming endogenous — the tooling generates scope. A CTO could equally read it as 27% more valuable work getting done, and that reading is not wrong. Either way it stopped being a variable both parties fully control at signature, which is a problem for a document whose entire function is to fix it.

Fixed Scope vs Flat Rate Software Development: The Rubric

Score your engagement on the five rows below. This is not a close call in either direction — most engagements land clearly on one side.

DimensionPoints to fixed scopePoints to flat rate
What you are buyingThe estimate itself is the product — a date, a number, a signed commitment you can hold someone toWorking software, evaluated continuously
Discovery remainingNear zero. The work is well-trodden and the spec is stableReal. Requirements will move as the thing gets built
CodebaseGreenfield, or a well-understood system with a known surfaceLegacy of unknown depth, or a mixed portfolio
Governing constraintProcurement mandates a scope document, or the scope is the compliance artefact (regulated audit trail)Commercial judgement, exercised by someone who can reprioritise
Change frequencyRare and consequential. You want every change to be a negotiationWeekly. The change-request overhead would exceed the changes

Three or more rows on the left and you should buy fixed scope. There is no argument in this article that beats a procurement rule or a compliance artefact, and pretending otherwise would be dishonest.

Three or more on the right and the fixed-scope contract is charging you for a forecast that the evidence above says nobody can currently produce.

What Holds Up in the Flat-Rate Model

Buy capacity, not scope — and be precise about the promise. It is a supply commitment: this much work in flight, continuously, reprioritised whenever you like. It is not a promise about output volume, and anyone selling it as one is selling the metric we are about to advise against.

Do not grade it on throughput. Output volume rose in every dataset above while delivery got worse. Thoughtworks has "measuring collaboration quality with coding agents" in Assess for exactly this reason: first-pass acceptance rate and iteration count, not units shipped.

Cap batch size contractually, not aspirationally. Review is the binding constraint, review time moved 441.5%, and agents make a 2,000-line diff trivial to produce and impossible to review honestly. That cap belongs in the agreement.

Own the spec — and notice what this costs the model we are arguing for. Under a fixed-scope SOW, the SOW is the specification. Drop fixed scope and you have removed the artefact that stated what was owed. Replace it with something you can point at, or you have bought a relationship rather than a deliverable.

Where the Flat-Rate Model Is Weakest

We sell flat-rate monthly development, so treat this as an interested party getting ahead of the obvious objections.

"You attack fixed-price for charging me for uncertainty I may not consume, and your alternative is a retainer I pay whether I use it or not." Correct, and the symmetry is real. The difference is not that one is risk-free. It is the repricing interval: fixed scope locks a guess for the length of a project, month-to-month re-decides every thirty days. With a base rate moving this fast, we would rather re-decide monthly. You are still paying for capacity you may not use.

Adverse selection, and we published the map. If greenfield work runs 35–40% better and legacy runs at 10% or less, a capacity model gives the vendor a quiet incentive to prefer the easy queue item. Fixed scope prices the ten-year-old billing module explicitly. Capacity does not price it at all. If you buy this way, that is the thing to write into the agreement and the thing to watch for.

"Cancel anytime" is an exit, not a remedy. It is not an SLA, and it is not a defined deliverable with a legal consequence attached. If you need those, fixed scope gives them to you and this does not.

None of this is new. It is a retainer with a work-in-progress limit. Design studios have sold the same structure since around 2018. The case for it is not novelty — it is that the stationarity assumption underneath fixed scope stopped holding.

The Uncomfortable Summary

If the estimate is the product, buy fixed scope. Those cases are real and the rubric above will find them.

If working software is the product, it is worth noticing what you are currently buying it with: an instrument whose central number moved 37 points in twelve months while the people who measure this professionally abandoned their own experiment.

See how Kuaray's flat-rate development works — one monthly rate, unlimited requests in the backlog, one to four in flight depending on the plan, month-to-month, and you keep the code. If your situation scores left on the rubric, tell us and we will say so — we would rather you got the diagnosis right than bought the wrong contract from us.

Share:LinkedInX