Beyond Quality
An open-source community for deeper enquiry.
The management-layer implementations
Part of Governance of quality.
What the standards leave to the organization
ISO/IEC 38500 is written as principles on purpose. In its own words, a principle-based standard identifies the outcomes of applying the principles without prescribing methods, structures, processes or techniques (5.1). ISO 37000 asks the same of the governing body’s policies: they give guidance on what is to be fulfilled rather than detail how (6.3.3.1.2 d). ISO/IEC 38507 keeps the split for AI: the choice of tools is a management decision, made within the guidance the governing body gives (4.2). The standards tell the governing body what to demand. How the organization meets the demand is left to the organization.
This research gives the bigger picture for quality: what the governing body states, and what comes back to it. The governing body writes one short document, the statement. It lists the qualities the company bets on, and each bet says the level to hold, what holding it is expected to earn or protect, what it costs, and when it is reviewed. It lists the minimums law and contract require, one common rule for every other need, and how much risk of each kind the company accepts (value.md). Management works inside those numbers and reports back in five kinds of reports (interface.md).
Economics of Testing is the “how” for risk management in QA, and it works without the governing body taking part. A QA lead can bring the managers together, collect what they are afraid of, turn the fears into a register of risks, and budget the register with them into a testing strategy, without asking anyone above. This research does not replace that. When the top is willing to take part, the top supplies two numbers the managers otherwise have to establish among themselves, what each level is worth to the company and how much risk the company accepts, and it receives their results as its reports. Nothing inside the economics research changes.
QA in the Age of AI-Accelerated Development clarifies how agents change what that work depends on: the knowledge people carry about how the code works and about why it exists. The erosion thesis rests on that.
Economics of Testing: risk governance, oversight and performance
The economics research treats testing as an investment: money spent now so that failures cost less later. It works out what the checks cost and what the failures they prevent would cost. Its method is a loop of four steps. QA convenes it, and the managers decide in it. The loop works on the loss side: what a failure would cost, and what it costs to prevent it. What a level of quality is worth to the company is not its subject; the governing body states that in the statement (value.md). When the governing body takes part, every risk in the register is a way of missing a level the statement sets, and the cost of the failure is the worth that level protects. The value is stated first, and the risks are derived from it.
Step 1, identify and quantify risks. QA brings together everyone who has something to fear from a failure, the managers, the engineers, product, operations, security and support, and asks what they are afraid of. Each fear is written down as a risk with its consequence: if the payment gateway times out under load, customers cannot pay and the sales are lost. Then the group puts numbers on each risk, always as three figures, the best case, the likely case and the worst case. Engineers estimate how many hours or days it would take to investigate and fix the failure and what that would do to delivery: one day if it goes well, two most likely, four if it does not. The managers estimate what the failure would cost the company: lost revenue, the cost of the incident, customers lost, fines, damage to reputation. Three things come out of the step: the list of risks, the ranges, and a written agreement on which risks matter most and why.
When the governing body takes part, the managers no longer estimate what a failure would cost. The statement already says what each level is worth. If the checkout path is held at four nines because customer success expects that level to keep three to five renewals a year, then a failure of the checkout path puts those renewals at risk, and customer success owns that number. The managers take the cost of the failure from the statement and spend the session on the question only engineering, QA and support can answer: how could we miss this level? A risk to a need that has no bet goes up as a proposal, and the top decides whether the need gets one (value.md, the needs without bets).
Step 2, categorize and prioritize. The managers decide together what matters most. They rank the risks by how likely each one is and how much it would cost, weighting money, safety, legal exposure and reputation the way the company weighs them. Then they draw the lines. Above one line a risk must be reduced. Below another the company can live with it. In between, a risk can be carried only with controls in place and a named owner accepting it in writing. The research says the lines should reflect the company’s risk tolerance. Today the managers have to agree on that tolerance among themselves, and the argument is often won simply by authority. The usual form is the product manager’s: this feature is urgent, so it ships now, tests or no tests. “Urgent” carries no number, QA’s “it is risky” carries none either, and two claims without numbers cannot be weighed against each other, so the senior side wins.
When the governing body takes part, the tolerance is not the managers’ to agree. The statement says how much risk of each kind the company accepts, and above what size of problem management must escalate instead of deciding alone. Where the damage would be irreversible, a security breach say, the appetite is near zero and anything close to it needs a sign-off. Where the damage is recoverable, a slow monthly report, the appetite is wide. The managers still do the ranking. How likely a failure is, and how badly it would hurt, are still their judgment and still argued over. What they no longer argue over is how much risk the company may carry, because the top has set that number. The product manager’s week of delay gets a number too: product with engineering estimate what postponing this feature costs, and what the failure QA fears would cost comes from what the level is worth (interface.md, the decision rule). The case is decided on the two numbers against the appetite, and neither side needs to outrank the other. Two reports come out of this step: how big each risk currently is, next to the appetite for its kind, and which risks were accepted, with the owner, the reason, the review date and the signature.
Step 3, decide where, how and how much to test. For each risk, QA picks the cheapest checks that give credible evidence. For a calculation error that might be a unit test; for a third-party API, a contract test; for a risk that only shows under real traffic, a canary rollout. Some risks need checks at more than one level. QA writes down, for each risk, which checks cover it, what evidence those checks produce, and who accepted whatever risk remains. That record is the test strategy. The managers pay for it out of their own budgets.
When the governing body takes part, QA still chooses the checks. The money is no longer the managers’ to find in their budgets: it was set when the bet was placed. The engineers who will hold the level estimated what holding it would cost, and that estimate is written in the bet as a range. The checks for the checkout path have to fit that range. The statement also says which needs get no testing money on purpose. The invoice export that runs once a month is kept as good as such exports usually are, so it gets one basic check and nobody asks for more. Where law or contract prescribes the checks, QA does not choose: aviation, automotive and medical software standards say which verification must be done and what evidence must be kept, and those checks are in the strategy whether or not anyone wants to pay for them. Nothing from this step goes up to the governing body. The test strategy stays with QA. When the top wants to know what stands behind a risk number, whoever does the assurance check (interface.md) reads that strategy.
Step 4, review and rebalance. The managers review the strategy on a schedule, once a quarter say, and whenever something big changes: a new regulation, an architecture change, a spike in defects reaching customers. They compare what the checks cost with what they found, and they shift money towards the checks that reduce the most exposure and away from the ones that found nothing worth their price. The research insists that the people who can decide are in that meeting; otherwise nothing gets rebalanced.
When the governing body takes part, the schedule is not the managers’ to set. Each bet carries its own review date, the numbers the departments already track are the baseline, and every trend is read against the appetite the top published. Most of what flows up comes out of this step. Whoever held the level reports what holding it actually cost, next to the estimate. The department that backed the bet reports what the level earned or protected. The managers report how each risk’s size has changed, the breaches and the warnings before them, and proposals to add a bet, drop one, or fund one differently when the strategy cannot hold a level for the money set aside.
Four of the five reports come out of the loop in full. The fifth, the bet check (each bet checked at its review date against what it was expected to earn and to cost), comes out in half: the loop supplies what holding the level cost, and the department that backed the bet supplies what it earned or protected.
In the standard’s terms, the loop, run inside the statement, gives the governing body what it is expected to demand. It gives the body a sound process for managing risks, in which the managers put the testing money where the risk is largest (38500 5.10.2). Under the statement the size of a risk comes from what the level at stake is worth, so the value is decided first and the risk follows from it. It keeps the appetite the body set both defined and managed (5.10.3). It keeps the body informed of material breaches and of patterns of breach (5.5.2). It measures performance against expectations the body stated, at intervals short enough for the parts of the system that change fast (7.2.6). The sign-offs from step 2 and the record from step 3 are what the body can point to when it has to show it was in control (7.2.7). They are also evidence the body can check for accuracy, as ISO 37000 asks it to (6.4.3.1 d). A proposal is what the standard expects when management finds a risk outside the appetite: management proposes a change to the numbers, and the body decides (38507 A.2).
The loop covers the risks to product quality. Cyber, financial and legal risks have their own management processes and their own standards. The hub’s scope leaves them out; the same interface would serve them.
Once the statement exists, the checks QA funds protect value as well as guarding against loss. A check that closes a risk protects the return the level was expected to bring, renewals kept or deals won, and the top can see the check and the return together. In the standard’s words, the bets are the company’s value generation objectives for quality, and the checks are the investment those objectives guide (38500 5.3.3).
AI-era testing: what the diagnosis means for governance and accountability
ISO/IEC 38507 asks two questions: how to keep governing when the organization starts using AI (4.2), and how to stay accountable when it does (4.3). It answers both with duties of the governing body. Governance that worked before may no longer fit once AI is in use. The body should treat its intended use of AI as part of its risk appetite. The body stays accountable for decisions made through AI and for the controls around it, and it cannot blame the AI system. The standard covers every use of AI in the organization, whatever form the AI takes (3.1.1). Its examples run from credit scoring to a grammar suggestion in a document (4.2, 5.4). The ai-era-testing research takes one such use, agents building and testing the organization’s software, and studies what that use does to the knowledge people carry about the software: how it works, and why it is the way it is.
The research starts from two symptoms (analysis). The first: testing after the code is written cannot keep up. Agents write code faster than testers can check it, and the queue of unchecked work grows. The second: testing after the fact can no longer test properly. To verify a feature, people need to understand two things, how its code was built and why it exists, and with agents writing the code nobody carries either:
- comprehension debt: nobody has a deep understanding of how the code works;
- intent debt: people lose the knowledge of why the code exists, what business problem it solves, and what trade-offs were accepted.
The first symptom is a cost, and the managers can count it in the four-step loop of Economics of Testing. Agents multiply the code a team produces, and the number of features with it. Every change still has to be checked after the fact, by a tester who tries it or by tests someone writes, and trying it or writing the tests takes a person as long as it always did. So the queue of unchecked changes grows faster than the testers can work it down, and to keep up by hand the testers would have to multiply faster than the code does (with agents).
Paid inspection tools do keep up with the volume. They are built for it. At one review tool’s published price, a team merging twenty changes a day pays about $10,000 a month for inspection alone, and more as agents write more (evaluation). What such a tool judges is the code alone. Quality is the degree to which the product satisfies the needs of its stakeholders (definitions.md), and the research starts from the same point: quality is not a property of code in isolation. Whether the code does the right thing depends on the business problem it solves, on what a failure would cost and on the trade-offs the team accepted, and a tool reading the code has none of that, so it catches the shallow faults and passes the clean code that solves the wrong problem. Letting the agent test what it built fails at the same point: the tests are generated from the code, they agree with it by construction, and a wrong implementation passes them (lifecycle drift). So the bill for checking after the fact grows with the code however it is paid: many more testers, or inspection that cannot judge whether the code does the right thing, or both. A company that pays for neither ships its changes unchecked and pays for the failures that reach customers instead, and what those failures cost is what the levels at stake are worth (value.md).
Before agents, the people who built and tested the software learned it by doing the work, and nobody paid for the learning (pre-AI baseline). Take the refund flow the research uses as its example: refunds over $500 need a manager’s approval, fraud cases skip the manager and go straight to legal, and partial refunds on subscription disputes are prorated to the day. Suppose a team had written it by hand. The developer who wrote the $500 rule understood the code, because they wrote it, and knew what the rule was for, because they had asked product while writing it. The tester who tested refunds learned where they went wrong and tested there harder the next time. Some of it they wrote down, when they judged a rule worth the writing, and the judgment of what was worth writing came from the same experience. What each of them learned went into the team’s shared understanding of the system, and that understanding held both kinds of knowledge, how the code worked and why it was there. With it the team could check any change on both counts: that the code does what the rule says, and that the rule is still the right one. Every step of the four-step loop of Economics of Testing ran on the same knowledge: what could fail and what it would cost in step 1, whether a check is credible evidence in step 3, whether the risks had changed in step 4. The people the managers asked knew, because they had built and tested the system. In the research’s words, “every step of this model depends on business domain understanding”.
For the company, handing a whole feature to an agent and taking back the build is much like outsourcing the feature to a contractor. The company knows what it asked for, it can check what came back against that, and it does not see how the work was done. Outsourcing is a risk the governing body already has a name for: oversight of sourcing arrangements can be a key governance issue (38500 5.5.2). Three things differ. A contractor has a name, a contract and a reputation. In the standard’s own comparison, a human can be questioned, double-checked, assessed by their standing and reputation and held accountable, and the code they wrote can be tested against what they said it does (38507 6.7.4). The agent has none of that, and the company alone answers for what the agent produced (4.3). A contractor can be asked afterwards why a rule is there. The agent keeps no understanding from one session to the next, so it cannot say.
The third difference is the research’s own point. A contractor learned the system while building it, so the knowledge exists in someone’s head, and the company can buy it back: a handover, documentation, a clause that keeps the people who know. With an agent nobody learned. The research calls this the comprehension inversion: “the code exists, but nobody has the deep understanding the author traditionally had” (with agents). With outsourcing the knowledge is in another company’s heads; with the agent it was never formed. And the company can check only what it asked for. A feature handed over in one prompt was asked for in one prompt. Nobody asked for the decisions the agent made on the way, so there is nothing to check them against.
The second symptom takes that knowledge away from the four-step loop of Economics of Testing: the managers can no longer say what a failure would cost in step 1, and QA can no longer say whether a check is credible evidence in step 3. The two debts are that loss, and they are not alike. Comprehension can be bought back. A tester who spends long enough with the refund code learns how it works; the research’s cost table lists understanding the code as the expensive step of any change, where before agents it cost nothing extra, because the person who wrote the code understood it. Intent cannot be read from the code at any price, because the why was never in it. A tester who knows every line of the refund code but not why fraud cases skip the manager cannot say what the rule protects, so they cannot say what its failure would cost.
The research works the difference between the two debts through the same flow, written by an agent (research question).
In the first case the team knows the three rules exactly, but nobody has read the code closely in the six months since it was written. A customer is refunded twice. The team turns on logging, compares what the code does with the rules it knows, and finds the fault. It takes longer than it would in code they had written themselves, but they fix it, and the product keeps doing the right thing meanwhile, because the team can tell right from wrong.
In the second case, two years later, the people who set the rules have left. The records of the A/B tests the team ran back then are gone, and the compliance document that set the $500 threshold has been superseded. The team now running the code has read every line and knows exactly what it does. What nobody knows is why. Why do fraud cases skip the manager: a compliance requirement, an operations decision, or a leftover from an A/B test that nobody cleaned up? Why is a partial refund prorated to the day and not the hour: does any customer still need that? So the team cannot say whether the code still does the right thing, cannot decide whether to keep, change or remove a rule, and cannot tell whether a new regulation fits what is already there. The research draws the line there: lost comprehension costs speed and risk, and it can be recovered; lost intent costs direction, and it compounds.
Read at the governance level, intent debt is the wearing away of the first of governance’s three components, direction. To direct is to communicate desired purposes and outcomes (38500 3.1). Management directs too, under authority delegated to it (3.1, Note 2). Somebody directed the refund rule. When nobody can say why, nobody can tell whether the system still does what was directed, or what to build next. The standard has its own example of the same gap: an instruction that is clear enough for a human operator and too imprecise for an AI system to execute correctly (38507 4.2, EXAMPLE 1). Among its sources of risk it lists unclear specifications, where the problem and its solution are no longer tightly bound (6.7.4).
Comprehension debt wears away the second component, oversight, which the standard splits into two tasks. To evaluate is to make informed judgements (38500 3.2). To monitor is to review as the basis for decisions and adjustments (3.8). Both need people who understand the system well enough to supervise it, step in when it goes wrong and change it. The standard has a term for how that understanding disappears: “agent atrophy”. People in roles where AI does the routine work see less and less of the work themselves, and the standard counts the hollowing out of their skill as a risk to the organization (38507 6.7.4). It lists the possible loss of organizational knowledge among the implications of using AI (4.2). It states that an AI system has no equivalent of human understanding of context (6.5). It warns that staff come to assume a sophisticated system is error-free, so the organization needs a process to validate the system’s output (6.6.1).
Accountability is the third component, and it does not wear away. It stays where it is. The governing body answers for decisions made through AI as it does for human ones, and it cannot delegate that (38500 6.1). The AI standard says the same: the body takes responsibility for the use of AI rather than attributing it to the AI system (38507 4.3). The body also has to explain how a responsibility was fulfilled (ISO 37000 3.2.2, Note 1). A body that has handed the writing of code to agents still owes that explanation, and the refund team above could not give it. That is why the research’s conditions concern the body.
Two things differ when the AI is a coding agent rather than a system deciding cases, and neither changes the body’s duties. The first is where the controls can act. A credit-scoring model decides every application afresh, so its mistakes are spread over thousands of decisions, and the organization can watch them: refusal rates by segment, defaults against predictions, drift over the months. The standard’s controls are written for that stream. The body monitors the decisions and output such a system generates and has management keep it within acceptable bounds (38507 6.3). It has the system reassessed when its behaviour changes over time (5.2.3). It asks for alerts that bring a person in when the numbers change faster than the reports come (38500 7.2.6).
A coding agent decides once. Take the refund flow again, and suppose the instruction to the agent said that refunds over $500 need a manager’s approval and said nothing about fraud cases. The agent decides on its own to send them to legal without approval. That decision was made once, while the code was written, and it is frozen into code that applies it to every refund, identically, for as long as the code runs. There is no stream of decisions to watch, because every run agrees with every other. Watching the product shows that it consistently does what it does; it cannot show whether that is what was directed. A wrong rule looks exactly like a right one, and it passes every test generated from the same code. It is found only when someone who knows what was intended compares intent with behaviour, usually when a new regulation arrives and a decision is needed. The one moment a person could have caught it was the moment it was decided. So the controls for this use of AI act while the code is built: on who decided each step, and on whether the evidence that the code does what was directed comes from somewhere other than the code.
The second difference is what the body can see of the use. A credit-scoring model is a named system with a vendor or a team and a place in the process, and a regulator asks about it. The body knows the model decides applications, and it knows the process around it, who reviews and at what threshold, because that process is the product. With agents, the body may well know they are used; many boards announce it. What it knows is the headline: agents build our software, this vendor, this many licences. What it does not know is how the code is built: what is handed to the agent in one piece and what in steps, who decides each step, whether the tests behind a risk number were written by the same agent as the code. Nothing in the product shows it. Agent-written code is indistinguishable from code a person wrote once it is merged, the repository does not record which lines an agent produced, and the product ships with no AI in it. The standard’s oversight conditions are about exactly the part the body cannot see: the people using the AI, what they understand, what they may decide, whether they can intervene (38507 6.2). The standard leaves the choice of tools to management, within guidance from the body (4.2). Here the body may have chosen the tool itself, and guidance on the practice was never given, because the practice looked like developers working faster. The standard lists the result among the failures it expects: a use of AI left outside the scope of the governance that exists (4.3). Until the body names the practice, how its software is built with agents, as something it governs, nothing in the oversight clause attaches to it.
One more difference the ai-era-testing research does not cover: the agent changes while the organization is using it, and not by the organization’s decision. The model behind it is updated by its vendor, the tool around the model changes its defaults, and the instruction that produced sound code in March produces different code in June, with no notice of what changed. A compiler changes with its versions too, but the same source has to mean the same thing under the new version, the organization can stay on the old one, and it can test its builds on the new one before switching. A coding agent gives none of that: the same instruction does not reliably produce the same code twice even within one version, the tools update themselves, and no specification says which properties of the output survive a change. The standard’s clause on systems that change over time is written for a product that retrains in operation; here it applies to the tool, and the assurance that it still works within acceptable bounds has to be repeated each time it changes (38507 5.2.3), which is often. So the body does not approve the way the company builds software with agents once. What holds still is the check on each output: a person deciding each step and judging its result does not depend on the agent staying the same.
The research puts four conditions on any answer (research question). Each one protects something the standards already require.
- Prevent intent debt from accumulating while the software is built. This is the standard’s direct task (38500 3.1). The AI standard adds that decisions made with AI stay aligned to the organization’s objectives, with a named person accountable who has the authority and the tools (38507 6.3).
- Keep comprehension sufficient for supervision, intervention and evolution. These are the evaluate and monitor tasks (3.2, 3.8). The AI standard requires of any human using or responsible for AI that they understand the system, can decide, and can intervene (38507 6.2). 38500 adds that a tool must not give its user more power than the user’s authority (7.2.5).
- Produce human-authored anchors for the lifecycle artifacts: docs, decision records, tests and onboarding material written by the people who decided, saying what a change was for and how it was built. The standard wants each decision and its explanation documented (7.2.4 d). It wants the body able to explain how responsibilities were fulfilled (ISO 37000 3.2.2, Note 1). The AI standard keeps the body accountable for an AI system to its decommissioning, including the knowledge the system contains (38507 4.3).
- Keep those anchors coherent as the system is maintained and extended and teams rotate. The standard asks for viability and performance over time (5.12). It asks for a regular assessment of the governance framework (7.2.1). The AI standard asks for further assessments as an AI system’s behaviour changes over time (38507 5.2.3).
The research sorts the industry’s responses into three directions (evaluation), which it numbers 1, 2 and 3. Direction 1 inspects the code better after it is written. Direction 2 improves the tests that AI generates after the code is written: better test generation, better ways of deciding what the right result is, mutation testing to catch regressions. Direction 3 changes how the code is built so that the debts do not form. The first two are appraisal in the economics research’s terms: checks after the fact. They help with the first symptom, the volume, and not with the second, because by the time they run the debts have already accrued. Only the third acts where the controls can act, while the code is built, and only it can meet the four conditions. The research advances one candidate for it, as a hypothesis with an experimental programme: two or more people building together in small steps while an agent writes the code (proposal).
The proposal gives the AI standard’s oversight clause its people. The clause requires of every human using or responsible for AI that they understand the system, can decide, and can intervene (38507 6.2). In a building session, a product person carries the why and has the authority to settle trade-offs on the spot. An engineer keeps understanding by deciding each step and judging each result. A QA engineer maps the work to the risk register before the session, which is the four-step loop of Economics of Testing applied to one feature, and raises edge cases during it: what if the user submits twice, what if the gateway times out. The size of a step is the check that the humans actually decided. A step is small enough when it fits one commit and the team can write the commit message by hand, saying why and how. When a team tells the agent to plan a whole feature and build it unattended, the session is delegation again: nobody decides the steps, understanding stops building, and the clause’s conditions are no longer met. The research states all this as a claim to be tested, and its experimental programme is on the proposal page.
The session gives the governing body more than the clause requires. The team writes its commit messages, intent notes and decision records by hand. Those records are the explanation the body owes when it is asked how a responsibility was fulfilled (ISO 37000 3.2.2, Note 1). They are also the documented decisions the standard asks for (38500 7.2.4 d). The refund team above could not explain the fraud rule. A team that wrote the rule down in the session can. The product person’s business rules and the QA engineer’s edge cases exist before the code does. So the evidence behind a risk size has a source other than the code, which is what the assurance check asks for. The step size keeps the agent presenting options rather than executing them, the boundary the standard tells the body to watch most closely (38507 5.1). A step small enough to decide is also a step the engineer learns from, so the skill to judge the next step is kept. When the steps grow, judging them takes understanding nobody is building, and accepting them takes nothing. So the steps grow further.
Five things the session does not do. It does not put the use of agents inside what the body governs, it does not fund itself, and it does not set the appetite for the knowledge risk. Those three are the body’s own acts. A product person and a QA engineer in the session cost money while the feature is built. Handing the whole task to the agent costs less at that moment, the vendor’s charge instead of two people’s time, and the build arrives sooner. Its cost comes later, as the debts, and no budget line shows it. The session holds the anchors together while it runs, and not across the years after it, when the team turns over and people who did not decide change the code. That is the research’s own open problem and its fourth condition, and the body needs it answered for viability over time (38500 5.12). It does not tell the body when the agent itself has changed; that takes the repeated assurance over the tool. It presupposes people who can still judge a step. A company that has already let that skill go has to rebuild it before the session can work. Last, it is a hypothesis with an experimental programme, so a body that requires the practice and pays for it is placing a bet, and the statement is where bets are written (value.md). The need is the company’s own: a system its people can still understand and change. The cost is the product person’s and the QA engineer’s time in the sessions. The return is what that understanding protects: failures caught before they reach customers, and the four-step loop of Economics of Testing kept running. At the review date the body puts the actual numbers next to the estimates and decides whether the practice stays. What the governing body has to do about it, to require the conditions and pay for them, is the subject of the erosion thesis and the practice page.
Where the two researches meet
The economics research has a rule against paying twice for one check: two checks that would fail for the same reason add little safety (step 3). A test generated from the code it tests is the extreme case. It comes from the same source as the code, so it passes while proving nothing, and a risk size built on such tests might be wrong while looking right. The ai-era-testing research describes the general form: tests and documents generated from the code agree with it by construction, so a wrong implementation is confirmed rather than caught (lifecycle drift). When agents write the code, whoever runs the top’s assurance check (interface.md) has one more question for each risk size: does the evidence behind it come from somewhere other than the code? ISO 37000 asks the body to assure itself that the reports and evidence it receives are accurate (6.4.3.1 d). The AI standard adds that AI monitoring other AI needs monitoring of its own (38507 6.6.2). The proposal’s session provides an independent source: the product person’s business rules and the QA engineer’s edge cases exist before the code does.
The debts have a cost, and the economics research has a category for it. The learning that came with the work before agents has to be paid for on purpose now: people writing down what a change was for, people building alongside the agent, people trained to supervise it. In the four categories of quality cost the economics research uses, that is prevention: money spent so that the problem never arises. The debts are what accrues when nobody pays it. The cost side of the erosion thesis starts here.
Agents are commonly used by handing over a whole task and receiving a whole build. That is the problem statement’s fourth observation seen from the economics side: everyone sees how fast the build arrived, and nobody counts the knowledge it did not produce. The proposal’s claim about step size is the countermeasure: keep each step small enough that people decide it, or the debts return. When the top takes part, the choice between paying for inspection after the fact and paying for people in the session is a decision like any other. Management works it out in money against what the level is worth and the published appetite, and the governing body sees both bills in the numbers that flow up.