When AI-written code causes an incident, “the model got it wrong” describes only part of what happened. Someone decided which systems the tool could reach, what context it received, how its output would be checked and who could authorize deployment.
Those decisions determine whether an error remains a rejected suggestion or becomes a production failure.
AI can improve development work. It can also increase the volume of changes arriving at a review process that was already stretched. The executive question is whether the complete delivery system can turn that extra output into reliable software at an acceptable cost.
The Evidence
A working demonstration is incomplete evidence
In a disposable prototype, a mistake may be cheap. The system has no customer data, no operational dependency and no important state to damage. Fast experimentation can be the right choice.
Production introduces different obligations. A change may affect access controls, billing, integrations or a process on which customers depend. The code needs to behave sensibly when inputs are malformed, a dependency fails or assumptions about user permissions are wrong.
A successful demonstration does not examine those conditions. Nor does a persuasive explanation from the agent that wrote the code.
NIST’s Secure Software Development Framework treats secure development as a set of practices integrated into the software lifecycle. That is the useful foundation here: security has to be produced by the surrounding process, whoever generated the first version of the code.
The Access
Decide what the agent may reach
A tool that proposes a change in an editor has a different risk profile from an agent that can execute commands, install dependencies and contact external systems. Adding those abilities changes the scope of the deployment decision.
Management should require a clear account of the data, tools and environments the agent can access. A task that requires reading part of a repository does not automatically justify access to production credentials or unrelated business systems.
The OWASP guidance on AI agent security recommends restricting tools and permissions to the task, validating sensitive actions and maintaining human controls. The business purpose of those boundaries is to keep a mistaken or manipulated instruction from producing an unrestricted consequence.
The agent should receive the relevant requirements and architectural constraints. It should also be prevented from quietly changing the controls that are supposed to constrain it. An instruction to preserve a security boundary has little force if the same agent can remove the check enforcing it.
The Verification
Faster generation requires credible verification
The quantity of code produced is a poor measure of engineering progress. Useful progress is a change that meets the requirement, fits the existing system and can be operated safely.
Automated checks help teams examine that evidence consistently. Their depth should match the change: relevant tests, dependency and secret checks, security analysis and performance examination where the risk warrants it. The checks also need the authority to stop delivery.
A green result deserves scrutiny if the same agent changed both the implementation and the test that supposedly proves it. The test may have been weakened, may reproduce the implementation’s mistaken assumption or may no longer exercise the real failure case.
OWASP’s secure coding guidance for AI-assisted work specifically addresses test deletion and weakening, changes outside the requested scope and human accountability. The practical consequence is that review must cover what changed in the verification itself.
More checking is not automatically better. The goal is credible evidence about the important failure modes. A large collection of superficial tests can consume time while leaving the central requirement unexamined.
The Review
Human review should follow the consequence of failure
A spelling correction and a change to customer-data isolation do not require the same depth of review. Treating them identically either wastes attention or gives the consequential change too little of it.
The reviewer needs enough context to understand the impact. They must see the actual change, relevant dependencies and the evidence from testing. Approval based only on the agent’s summary leaves the producer of the change in control of what the reviewer notices.
The business should define which classes of change require specialist involvement and which release decisions must remain with an authorized person. The responsible person may use AI to assist the review, but another generated opinion is not, by itself, independent assurance.
THE THROUGH-LINE
Review capacity belongs in delivery planning. If generation becomes much faster while review capacity stays fixed, the organization can accumulate a larger queue of unverified work. Calling that queue productivity disguises an operating constraint.
The Recovery
Recovery is part of readiness
Even a disciplined team can ship a defect. Readiness therefore includes the ability to detect the problem, limit its impact and recover.
That requires an appropriate release approach, useful monitoring, a tested recovery method and someone who can act. A backup that nobody has restored is weak evidence of recoverability. A rollback plan may also be incomplete if the change has altered data or triggered external actions that reversing the code cannot undo.
The record of the change should connect the requirement, the modified components, the checks, the approval and the deployed version. Preserve the evidence needed to investigate without unnecessarily retaining secrets or sensitive user data.
These are operating costs. Include them when comparing faster development with the work required to keep the resulting system reliable.
The Call
Grant more autonomy after the boundary has been tested
Start with a scope the team can supervise well. Broader authority should follow evidence that the agent handles the task within its limits, that failures are detected and that recovery works.
A run of successful ordinary cases does not prove that an agent respects a boundary. Test cases it must refuse, incomplete information and unexpected dependencies. Check that it cannot bypass controls when the easiest route to completing the task would be to weaken them.
Before expanding access, ask the responsible lead to explain what the agent can change, which failure blocks delivery, who approves consequential changes and how a bad release is contained.
The useful promise of AI coding is better software delivered with less wasted effort. To judge whether that promise is being met, measure verified delivery and operating outcomes—including review, rework and incidents. The number of changes generated is only the workload entering that system.