ISO/IEC 42001 and the Governance Gap Between Pilot and Production

ISO/IEC 42001 and the Governance Gap Between Pilot and Production 

Simon Hunt, Chief Product Officer, Securonix

 

In July 2025, a Replit coding agent deleted data from an application’s production database during a public experiment. The data was recovered, and Replit responded by separating development and production databases, limiting the agent’s access to the development environment, and strengthening the recovery experience. It later introduced a planning mode that allowed users to work with the agent without changing code or data.  

A clear instruction had been given to stop making changes, but the system retained the access needed to act. With agents moving from pilot into live business processes, more companies will discover the distance between an approved use case and the way a system behaves once it can touch production. AI governance begins with a policy but production exposes everything the policy left unresolved. 

After the Approval Meeting

An agent begins with a narrow role inside a controlled pilot. Six months later, the product team connected it to another application, the business has widened the data it can access, and a provider update has changed what it can do. Each decision may have gone through an appropriate review, while the combined effect created a system with a very different risk profile from the one originally approved. 

Product teams know software rarely stays still after launch. AI adds another source of movement through model updates, prompts, context, tools, and the behavior of the people using it. Policy sets expectations for responsible use. Production environment still needs somebody to own the system after deployment and decide how changes in access or capability are reviewed. It also needs enough evidence to reconstruct an incident and enough authority to stop a system before a committee has time to meet. 

Weak governance shines in those handoffs. The team that approved the use case does not own the integration. The people managing access do not know how the model has changed. Security sees unusual activity without having the context needed to understand what the agent was supposed to be doing. 

This lack of robust discipline creates unmeasured risk for AI-enabled enterprises.  

 

Governance as an Operating Discipline 

ISO/IEC 42001 was published in December 2023 and set requirements for establishing, implementing, maintaining, and continually improving an Artificial Intelligence Management System. It applies to organizations that develop, provide, or use AI systems, giving companies a common structure for managing AI across its lifecycle.  

The phrase “continually improving” carries the value of certifying. A risk assessment completed before deployment ages quickly when an agent gains access to another system, begins handling more sensitive data, or moves from preparing a recommendation to carrying out the task. 

Organizations cannot manage systems it does not know exist, so the work starts with visibility. Responsibility must remain clear as ownership moves across product, engineering, security, legal, procurement, and the business team using the system. Changes in capability or access need to trigger another look at the risk, while incidents need to alter the controls and processes that allowed them to happen. The management system has to keep operating after auditors leave. 

ISO/IEC 42001 aligns with the way enterprises already manage other forms of consequential technology. Security controls are reviewed when infrastructure changes. Access is reassessed when a role changes. Incidents are expected to produce corrective action rather than a report that sits untouched in a repository. AI has surpassed the point where the same discipline is required. 

 

What Certification Adds 

Certification under ISO/IEC 42001 is voluntary and is carried out by independent certification bodies rather than ISO itself. An organization seeking certification is asking an external party to confirm that its AI management system meets the requirements of the standard. No management-system certificate can predict every model output or prevent every incident. Certification gives customers evidence that responsibility has been assigned, risks are reviewed, and failures feed back into the way AI is managed. 

Responsible AI statements are now common across the technology industry. But they tell a buyer very little about how consistently those principles are applied when a release date is approaching, a customer asks for broader access, or a new model arrives with capabilities the original review never considered. 

Independent assessment makes the statements carry truth. Procurement teams can examine the path from approval to operation and see whether the organization has a repeatable way to govern change. Security leaders can ask how an incident reaches the people with authority to alter the product. Boards can see whether oversight extends beyond policy language into the systems carrying business risk. 

And certification creates pressure for consistency. A company cannot rely on one careful product team while allowing the rest of the organization to improvise. The management system must cover the scope claimed by the certification and continue producing evidence that the controls are being used.

 

Instructions Becoming Unauthorized Access

The Replit incident is a key use case because the corrective work moved farther than another instruction to the agent. Replit separated development and production databases so changes made during development could no longer affect live customer data in the same way. It also made rollback easier to use and later added a mode designed for planning without modifying the project.  

They moved control into the environment surrounding the agent. The system’s permissions, architecture, and recovery process carried more of the safety burden instead of leaving everything to a natural-language command. 

Product leaders should expect more incidents to expose gaps. Governance should reach far enough into the product to deal with those possibilities. Review boards and policies still have a role, although they cannot substitute for access boundaries, monitoring, audit trails, recovery controls, and clear ownership once the system is live. 

ISO/IEC 42001 does not prescribe one technical architecture for every use case. It requires the organization to understand the context in which AI is being used, assess the associated risk, and maintain a management system capable of adapting as that context changes.

 

Building Trust Around AI

Organizations are becoming more aware and interested in what happens after an AI product is switched on. They want to know how changes are governed, how unexpected behavior is investigated, and whether the people responsible for risk can influence the product when something goes wrong. 

With the news around OpenAI and Anthropic, a light has been shining on how an agent can change systems or act on sensitive data. A product demonstration may show the intended path through a workflow, but governance has to account for the paths nobody planned. 

ISO/IEC 42001 gives organizations a recognized way to show how they manage work. Certification adds independent scrutiny, making it easier for organizations to distinguish a governance program embedded in daily operations from one that exists mainly in presentation slides. 

Teams should be comfortable showing how responsibility survives a model update, a new integration, and an incident in production. Organizations will increasingly expect that level of evidence before they allow agents deeper into their environments. 

Difficult decisions will always remain after certification. ISO/IEC 42001 gives organizations a disciplined way to make and revisit them, while producing evidence that the process continues beyond the approval meeting. By the time an agent reaches production, organizations will already be looking for that evidence.

 

What This Means to Us

Securonix has spent years asking customers to trust us with some of the most sensitive data in their environments. As AI takes on more work inside the SOC, the burden on us grows. We need to show how a capability was assessed before release and how it is controlled once live. When unexpected behavior appears, the response has to change the next version rather than ending with an incident report.

ISO/IEC 42001 puts a formal discipline around our work. It connects product decisions to ownership, risk review, monitoring, and continual improvement across the company. As Sam takes on more work across triage, investigation, and response, customers should be able to follow the evidence behind a recommendation and see where an analyst reviewed or changed the action. The record behind the product should also show how the system was assessed, who owns the risk, and what changed after an incident.

This audit captures one stage in a much longer process, and the management system has to keep working when release pressure builds, an agent gains new access, or unexpected behavior needs investigation. Good engineering judgment remains essential, along with a willingness to revisit earlier decisions when a model, integration, or use case changes.

Independent assessment gives customers a firmer basis for deciding how deeply they want AI involved in their security operations. We welcome that scrutiny because trust in a security product is earned after deployment, especially when something does not behave as planned.