When AI writes code, who owns the mistakes?

AI-generated code accountability

Most enterprises can name the person who signed off on human-written code, but when you ask who’s accountable for code AI agents worked on, the answer takes much longer. In February, Moltbook, the predominantly vibe-coded social network for AI agents, highlighted this issue perfectly. Its database was misconfigured and left open, exposing the authentication token for every agent on the platform along with 35,000 email addresses. Behind the platform’s 1.5 million agents were roughly 17,000 people with around 88 agents each. Nobody could track which agent had done what, or on whose instruction – making it impossible to identify who was responsible for compromising the database. Moltbook was, in the words of its founder, ‘a weird experiment’, so it’s easy to initially write this incident off as a cautionary tale for startups. However, the ratio of agents to people it ended up with is similar to what enterprises are now deliberately building towards.

This direction of travel is already measurable. AI now generates or assists in generating 61% of the average enterprise codebase, and 64% of engineering organisations describe the technology as widely adopted or fully integrated. Our 2026 State of Code Abundance report, a survey of more than 200 enterprise technology leaders, found that 52% reported a significant increase in production output directly tied to agentic coding. While these AI strategies are delivering the production benefits leadership is targeting, we’re starting to see the problems that inevitably arise from failing to maintain the same levels of accountability that existed prior to implementation.

Inconsistent enforcement

On paper, most enterprises have a sufficient review process for AI-generated code, but this is not enforced consistently. Of the leaders we surveyed, 93% declared that they have a formal process for reviewing and releasing AI-generated code into production; however, only 56% said that this process is always enforced. This gap in enforcement was survivable when hundreds of changes a week took place. Thanks to AI, this has scaled up to tens of thousands, which increases the quantity of unreviewed code reaching customers.

The main pressure in software development has now moved downstream. Only 35% of enterprise technology leaders name writing code as their primary delivery bottleneck, while 57% point to reviewing, testing and deploying what has already been written as the main barrier. Furthermore, 70% say that maintaining their automated tests is a bigger burden than writing code. Verification has become the time-consuming and expensive half of the job, but it is what decides whether an organisation can stand behind what it ships.

Responsibility is defaulting upwards

Ask organisations who is accountable when AI-generated code causes a problem in production and 46% point to the CTO or VP of Engineering. On the surface, that might look like clarity. However, this is because only 12% have a dedicated AI governance team, meaning in most organisations no decision was made to assign accountability, so it migrated upwards. The issue with this arrangement is that a CTO or VP of Engineering usually sits several layers away from the change itself, the agent that produced it and the review that let it through. This means that the people named as accountable aren’t close enough to the work to truly own any mistakes.

The consequences of this assignment of accountability are apparent, with 81% of organisations reporting an increase in production issues they can attribute directly to AI-generated code. When one of those issues reaches a customer, somebody has to reconstruct how the change got to production, who approved it and what was checked. In most organisations today, that reconstruction is mostly guesswork.

A question with a deadline

From the 11th of September, this reconstruction became a legal requirement for anyone selling software or connected products into the EU. Under Article 14 of the EU Cyber Resilience Act, a manufacturer placing a product with digital elements on the EU market who becomes aware of an actively exploited vulnerability has 24 hours to file an early warning with ENISA and its national computer security incident response team, 72 hours for a fuller notification and 14 days for a final report. This duty extends to products already on the market, not just what ships from now.

Twenty-four hours is enough time to complete the form, but working out which change introduced the flaw, which agent produced it and who approved it takes considerably longer. Financial services work to an even shorter deadline. TheDigital Operational Resilience Act gives a financial entity four hours from classifying an incident as major to notifying its regulator, with a 24-hour outer limit from the point of awareness. NIS2 applies a 24-hour early warning across energy, transport, health, water and digital infrastructure, and extends to manufacturing, chemicals, food and research. Every one of these deadlines assumes an organisation can rapidly reconstruct its approval process.

What has to change

None of this is an argument against agentic coding, as the output gains are real and most organisations have already greatly benefited from them. Slowing adoption to restore human review at historical ratios would forfeit those gains and leave the underlying problem in place. What desperately needs to change is enforcement. The processes to ensure this already exist, but need to be applied to every change, whether a person or an agent made it. Importantly, they need to be integrated into the pipeline rather than just existing in a policy document that gets consulted after something has gone wrong.

Right now, there are three simple questions that can be put to engineering leadership. Which of your controls apply to every change, rather than most of them? Who, by name, owns a failure in AI-generated code? Could you reconstruct, for an auditor or a regulator, how any given change reached production? Organisations that can answer all of these will keep the speed they have gained. Those that can’t are relying on nothing going wrong.

Loreli Cadapan

Loreli Cadapan is VP of Product at CloudBees. A product leader with over 20 years of experience building and scaling enterprise software, developer tools, SaaS, cybersecurity, DevOps, and DevSecOps, Loreli turns complex technologies and customer problems into products that deliver meaningful business value.

Author

Scroll to Top

SUBSCRIBE

SUBSCRIBE