Certified Open Mercato Agency · AI-native enterprise systems for manufacturers · Book a founder call
Blog/Collaboration
Collaboration

Working With a Software Agency: Who Does What, and When

Hero illustration for Working With a Software Agency: Who Does What, and When

A project can go wrong while the agency does everything it promised. The code gets delivered, the demos happen, the invoices match the estimate, and the thing still lands months late with half the value the business expected. When we look back at those projects, the missing piece is rarely on the vendor's side of the table. It is a decision nobody had the authority to make, a set of acceptance criteria nobody wrote down, or a benefit nobody measured after go-live.

We are the.good.code, a software agency working with manufacturers and operations-led companies across Europe. This guide covers the division of labour in a custom software project. It sets out what the agency owns, what stays with you, and which parts of your side of the work belong in the contract in writing.

Why do software projects overrun even when the vendor delivers?

Because the risk is not evenly spread, and the part of it you control sits outside the vendor's scope.

The best evidence on overruns comes from Bent Flyvbjerg's team at Oxford, who analysed 5,392 IT projects and found that cost overruns follow a power law instead of a normal distribution. Most projects land near their estimate. A minority go catastrophically wrong: 18% of IT projects exceed their budget by more than 50%, and within that group the average overrun reaches 447% [1]. Averages hide this completely. You are either in the ordinary tail or in the disaster tail, and what separates them is decided early.

Deloitte's Global Outsourcing Survey 2024, based on responses from more than 500 executives, points at the buyer's side of that decision. The most commonly reported shortcoming of outsourcing arrangements was the lack of benefit realisation tracking and reporting, meaning nobody on the client side systematically measured whether the promised value arrived [2]. The same survey found that 70% of executives say their vendor management function is not fully mature.

That is the honest framing of this article. Your agency can be excellent and your project can still fail, because a large share of the work belongs to you.

Who does what, and when?

Here is the split we use with clients. Nothing in the client column can be handed to the vendor without losing something.

  • Discovery — The agency owns: technical assessment, architecture options, estimates, risk register. You own: business goals with numbers attached, access to the people who do the work, the systems map, the decision on scope.
  • Design — The agency owns: wireframes, technical design, data model, integration approach. You own: brand assets and design constraints, feedback within an agreed window, sign-off on what the software must do.
  • Build — The agency owns: delivery, code quality, testing, security, progress transparency. You own: one decision maker with a mandate, answers to blocking questions, prioritisation when trade-offs appear.
  • Acceptance — The agency owns: working demo, release notes, defect fixes, documentation. You own: reserved time for user acceptance testing, real users doing the testing, written acceptance against agreed criteria.
  • After go-live — The agency owns: maintenance, monitoring, evolution of the platform. You own: adoption inside the business, measurement of the promised benefit, a named owner of the system.

The decision maker during build and the acceptance testing before release cause most of the damage when they are left vague. Both are unglamorous and both are yours.

Four client-side responsibilities worth writing into the contract

Contracts usually describe the vendor's obligations in detail and the client's obligations in one line about providing timely feedback. These four are specific enough to be written down, and specific enough to be missed.

1. One decision maker with a real mandate

Name one person and give them authority to approve scope changes up to an agreed value without escalating. A committee in this seat slows every decision to the pace of its next meeting. Projects slow down in the gap between a question being asked and an answer being possible, and that gap is a governance problem before it is a technical one. When the decision maker has to reconvene four stakeholders for every trade-off, the delivery rhythm collapses to the speed of the slowest calendar.

Write into the contract who that person is, what they can approve alone, and how quickly the project can expect an answer. Two working days is a reasonable target for anything that blocks development.

2. Acceptance criteria written in business language

Every task needs a description of what "done" looks like before anyone writes code, phrased as behaviour. When a dealer submits a configuration above the credit limit, the system holds the order and notifies the account manager. That sentence is testable by someone who has never seen the codebase.

Requirements management is where projects quietly lose their goals. In PMI's research on requirements as a core competency, 47% of unsuccessful projects missed their goals because of inaccurate requirements management [3]. Treat that as a survey of practitioners and not a measurement of projects, but the direction matches what we see: the projects that go wrong are usually the ones where nobody could say precisely what the software was supposed to do.

3. User acceptance testing as reserved time, not a good intention

You are the only party that can confirm the software does what the business needs. The vendor can prove that the code matches the specification. Only your users can prove that the specification matched reality.

The clearest public illustration comes from a US government audit. When the Department of Education rebuilt the federal student aid application, the agency authorised system acceptance testing to begin even though 26 of the 48 readiness indicators were incomplete. By March 2024, 55 defects had been documented after launch, seven of them critical and unresolved, and one of those miscalculated aid eligibility by ignoring family assets [4]. The acceptance gate opened early, and the defects surfaced in front of applicants instead of testers.

The practical version for a mid-sized company. Book the people by name and the days in the calendar. Give them scripted scenarios drawn from the acceptance criteria. Make their sign-off the release condition. Testing that happens in the gaps between someone's normal job finds the obvious defects and misses the expensive ones.

4. Benefit tracking after go-live

Decide before the project starts which numbers should move, who reads them, and when. Order processing time, quote turnaround, error rates, the hours a team spends re-typing data between systems. Take a baseline before the build, because after go-live nobody can reconstruct it with any accuracy.

This is the responsibility that Deloitte's respondents most often admitted to skipping [2], and it is the one that determines whether your next project gets funded.

What rhythm keeps a project honest?

A demo of working software on a fixed cadence, and a written change process. Those two habits do more for delivery than any reporting template.

Status reports describe progress. Working software demonstrates it. The difference matters because a report can stay green for weeks while the underlying work drifts, and a demo cannot. DORA's research programme has consistently found that small batches of change and fast feedback loops correlate with better delivery outcomes [5], which is the engineering version of the same idea: shorten the distance between a decision and evidence that it was right.

The change process is the other half. Scope changes are normal and healthy, and uncontrolled ones are what turn a project into an overrun. PMI's Pulse of the Profession found in 2018 that 52% of projects experienced scope creep [6], and by its 2021 edition the figure had fallen to 34% [7]. Both numbers describe the same mechanism. Changes with no route to be requested and estimated arrive informally, then get absorbed in silence until the timeline breaks.

Does changing your mind late cost ten times more?

Less than the industry folklore claims, and that is good news for how you plan.

The familiar cost of change curve says that a defect found after release costs a hundred times what it would have cost during requirements. It comes from work published in 1981, and the underlying data was never released for analysis. Researchers who went looking found that almost every citation of the effect traces back to that single source [8]. A widely reproduced version of the chart is attributed to an "IBM Systems Sciences Institute", which turns out to have been an internal training programme instead of a research body [9].

When the question was tested again on modern data, the effect largely disappeared. Menzies and colleagues examined 171 commercial projects delivered between 2006 and 2014 and found no evidence for the delayed issue effect. The ratio between fixing an issue early and fixing it during system test was 1.11, and the worst case they observed was a factor of three [8]. The original authors of the classic curve had already narrowed their own claim years earlier, putting the hundred-to-one ratio in large systems and something closer to five-to-one in small, non-critical ones [10].

The steepness of that curve is a variable you control. Iterative delivery, small batches, automated tests and short feedback loops flatten it. Long gaps between demos, unclear acceptance criteria and a decision maker who is hard to reach make it steep again. Read it as a property of your process instead of a law of nature, and the planning question changes from "how do we avoid all late changes" to "how cheap can we make them".

For high-assurance systems the old warning still holds. NASA's analysis of large aerospace programmes reported error costs escalating from one unit at requirements to between 21 and 78 units at integration and test [11]. If you are building something where a defect is dangerous instead of annoying, plan for the steep curve.

What has AI changed about your side of the work?

The volume of decisions arriving at your desk went up. The time available to make them did not.

Start with what the data supports. In DORA's 2025 research, based on nearly 5,000 technology professionals, 90% report using AI at work and more than 80% believe it has increased their productivity, while 30% report little or no trust in the code it generates [5]. Telemetry from delivery analytics vendors shows what that does to the throughput of work needing human approval: Faros AI reported teams with high AI adoption merging 98% more pull requests alongside a 91% increase in median review time [12], and its 2026 analysis of 22,000 developers reported a 441% increase in median review time [13]. Both are vendor-reported, so treat the magnitudes with caution and the direction as consistent with everything else. GitClear's analysis of 211 million changed lines found copied code rising from 8.3% to 12.3% of lines between 2020 and 2024 while refactored lines fell from 24.1% to 9.5%, also vendor-reported [14]. In Stack Overflow's 2025 survey of more than 49,000 developers, the top frustration with AI tools, cited by 66%, was output that is almost right but not quite [15].

Now the honest caveat, because the counter-evidence is strong. In a randomised controlled trial published by METR in 2025, experienced open source developers working in large, mature repositories took 19% longer with AI assistance, while believing they had been 20% faster [16]. That is the context closest to working on an established business system, and it argues against any claim that AI simply compresses delivery timelines. METR's own follow-up in February 2026 reported results too uncertain to call in either direction [17].

So we make the narrow claim. Whether the code arrives faster is contested. That more artefacts arrive per unit of time for a human to review, approve and correct is not. On our own projects the practical consequence is the client's: the window in which a misunderstanding can surface has become shorter. A twelve-month project offered dozens of natural moments to notice that a requirement had been read differently by both sides. A six-week build offers a handful, and each one you skip costs more, because the team builds further on the assumption in the meantime.

The response is more precision at the start. Acceptance criteria written before the work begins, a decision maker who answers in days, and demos frequent enough that a misreading surfaces while it is still cheap.

How we run this

At the.good.code we run this rhythm inside our development packages. That is one of several ways we work with clients, alongside Technology Strategy Advisory and full Platform Ownership and Development. A dedicated team, a shared task board, an estimate and acceptance criteria on every task before any code gets written, and a working call in a cadence set by the size of the package, usually two-week sprints, where the client watches the software run. Speed is why it matters. A six-month project leaves plenty of moments to catch a misunderstanding. A six-week one leaves a handful. What clients tell us afterwards is that they changed too. After a few cycles they know which requests are cheap and which are architecturally expensive. They raise them earlier. When something breaks, they already understand where it came from.

Every task moves through the same states, and the state names are the governance: nothing reaches development without an estimate and acceptance criteria, and nothing closes without review. Hours are tracked per task and reported monthly against what was booked. A client who reads that report can see where the money went at the level of individual work items.

When should you not hire an agency?

Three situations, and we say so in the first conversation.

The software is your competitive advantage and it changes weekly. Products at the core of how you win are better owned in-house, where the feedback loop between a customer conversation and a release has no contract in the middle.

You already have a capable team and a clear roadmap. Then you are buying capacity, senior review or a second opinion on architecture. That is a different purchase from a delivery partner.

Nobody can be released to run the project on your side. If the four responsibilities above have no owner with time in their week, the project will drift regardless of who builds it. This is the honest test we apply before quoting, and it is the reason we sometimes recommend waiting a quarter.

If you are earlier in the process and still choosing a vendor, we wrote separately about what to prepare before you shortlist anyone.

Frequently Asked Questions

How much of a software project is the client's work?

Enough that it needs a name in the plan. The recurring commitments are a decision maker available within days, feedback on demos in a fixed cadence, and a block of reserved time for user acceptance testing before release. We have found no credible published benchmark for the hours involved, so agree the shape with your vendor at kick-off and put it in the schedule instead of assuming it will fit around normal duties.

Who should write the acceptance criteria?

The business, with help from the vendor on phrasing. The person who will judge whether the software works should describe what working means, in the language of the process instead of the system. A vendor writing its own acceptance criteria is marking its own homework.

Can we delegate user acceptance testing to the agency?

The agency can prepare test scripts, environments and data, and it can run its own quality assurance. The judgement about whether the software matches how the business operates has to come from your users. When acceptance gates open before that judgement exists, defects surface in front of customers. The 2024 audit of the US federal student aid system documents exactly that sequence [4].

How should change requests be handled?

In writing, with an estimate, before the work starts. A change request should say what changes and what it costs in time and money, then name the work it displaces. Visible changes stay cheap. Uncontrolled scope change remains one of the most commonly reported causes of overrun [6][7].

What should we ask an agency about its use of AI?

Ask how AI-generated code is reviewed, what static analysis runs in the pipeline, and how regression tests are maintained. Then judge delivery by working software at an agreed cadence instead of by commit volume. With 30% of developers reporting little or no trust in AI-generated code [5], a vendor with a clear answer here is telling you something useful about its engineering discipline.

What to agree before you sign

The division of labour in a software project deserves more attention than it usually gets after the contract is signed. Name the decision maker, write the acceptance criteria in business language, reserve the testing time, and agree what you will measure after go-live. Those four commitments cost nothing and they separate the projects that land from the ones that end up in the expensive tail.

If you want to sanity-check the split before you sign with anyone, we run a free 30 minute conversation with founders under our Technology Strategy Advisory service. Bring your project outline and we will tell you which parts of the work will be yours.

Sources

  1. Flyvbjerg, B. et al., "The Empirical Reality of IT Project Cost Overruns", Journal of Management Information Systems, 2022 (n=5,392)
  2. Deloitte, Global Outsourcing Survey 2024 (more than 500 executives)
  3. Project Management Institute, "Requirements Management: A Core Competency for Project and Program Success", 2014
  4. US Government Accountability Office, GAO-24-107783, September 2024
  5. DORA / Google Cloud, State of AI-assisted Software Development, September 2025 (nearly 5,000 respondents)
  6. Project Management Institute, Pulse of the Profession 2018
  7. Project Management Institute, Pulse of the Profession 2021: Beyond Agility
  8. Menzies, T., Nichols, W., Shull, F., Layman, L., "Are Delayed Issues Harder to Resolve? Revisiting Cost-to-Fix of Defects throughout the Lifecycle", Empirical Software Engineering 22(4), 2017
  9. The Register, "Cost of fixing bugs: the numbers behind the folklore", July 2021
  10. Boehm, B., Basili, V., "Software Defect Reduction Top 10 List", IEEE Computer 34(1), 2001
  11. Haskins, B., Stecklein, J. et al., "Error Cost Escalation Through the Project Life Cycle", INCOSE International Symposium 2004 / NASA
  12. Faros AI, analysis of delivery telemetry, 2025 (vendor-reported)
  13. Faros AI, "The Acceleration Whiplash", 2026, 22,000 developers (vendor-reported)
  14. GitClear, AI Assistant Code Quality Research, 2025, 211 million changed lines (vendor-reported)
  15. Stack Overflow Developer Survey 2025 (more than 49,000 respondents)
  16. METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity", July 2025
  17. METR, "We are Changing our Developer Productivity Experiment Design", February 2026
Interested in working together?

A 30-minute conversation with the founders.

Start a conversation →

work.with.us;

Bring one real problem; leave with a clear answer.