CEO Corner: AI Operations: What Happens After the Model Goes Live? by Mark Hewitt
For much of the conversation around enterprise AI, deployment is treated as the finish line. A model is selected, an application is built, testing is completed, and the solution goes into production. The project team celebrates, the business begins using it, and attention naturally shifts to the next opportunity. But with AI, going live is not the end of the work. In many ways, it is where a different kind of work begins.
Traditional software is certainly not static, but AI systems introduce another level of variability. Models change, data changes, user behavior changes, business conditions change., costs change, and the quality of outputs can change without a single line of application code changing. As organizations move from copilots toward agents and increasingly autonomous workflows, those operational considerations become even more important. That is why AI Operations, or what we think of as Day 2 Operations, is an essential part of building an AI-native enterprise.
Production Changes the Question
Before an AI capability reaches production, the central question is usually straightforward: Does it work? Once it is operating inside the enterprise, the questions become more complicated:
Is it still producing the quality of output we expected?
Are employees actually using it? Where are humans intervening?
Is the underlying data changing? What does each transaction cost?
Are response times acceptable? Are security and governance controls working as designed?
Most importantly, is the solution continuing to create the business value that justified the investment in the first place?
These questions cannot be answered once and filed away. They need to be answered continuously. An AI capability that performed extremely well during a pilot may behave differently when thousands of people begin using it in ways the original team never anticipated. A workflow that was economical at a few hundred transactions per day may look very different at enterprise scale. Or, a model that was appropriate six months ago may no longer be the best choice today. Production gives us something experimentation cannot: evidence from the real world. The organization needs an operating model capable of learning from that evidence.
Observability Has to Evolve
Enterprise technology teams already understand monitoring and observability. We monitor availability, latency, infrastructure, errors, security events, and application performance. AI expands what needs to be observed. Availability still matters, but so does the quality of the response. Accuracy matters. Hallucinations matter. Token consumption and model latency matter. Retrieval quality matters. User feedback matters. Human corrections and escalations matter.
With agentic systems, observability becomes even more important. If an agent can take an action, invoke another system, access enterprise data, or make a decision inside a workflow, organizations need visibility into what happened, why it happened, what information was used, and where human oversight was required. This is one of the important differences between experimenting with AI and operating AI as an enterprise capability. We need to understand not only whether the system is running, but whether it is behaving as intended and producing the outcome we expected.
Model Performance Can Change Even When the Application Does Not
AI systems also introduce a reality that traditional application teams may not always encounter in the same way: the application can remain unchanged while its performance changes. The data feeding the system may shift. Retrieval sources may become outdated. User behavior may evolve. A model provider may release a new version. Business terminology may change. New products, policies, customers, regulations, or operating conditions may emerge.
The result is that an AI capability that worked exceptionally well when it was deployed can gradually become less effective. This is why monitoring model and data health needs to become part of normal enterprise operations. Organizations need mechanisms for detecting degradation, understanding its cause, and deciding whether the right response is a new model, different data, improved retrieval, a workflow change, additional human oversight, or something else entirely. AI Operations should not simply detect failure. It should create the information needed to continuously improve the system.
The Economics Need to Be Operated Too
There is another dimension of AI Operations that I believe will receive much more attention as enterprise adoption scales: cost. During experimentation, model costs can appear relatively insignificant. At enterprise scale, the economics can change quickly. Organizations need visibility into model usage, token consumption, infrastructure, data movement, licensing, retrieval, monitoring, security, and the other components required to operate AI reliably. Those costs then need to be connected back to the value the capability is producing. This creates something similar to a FinOps discipline for AI.
The most powerful model is not necessarily the right model for every workload. Some use cases may benefit from smaller models. Others may benefit from localized models, caching, improved retrieval, model routing, traditional automation, or combinations of technologies. The objective should not be to use the most sophisticated technology available. It should be to create the best-performing and most economical solution for the business problem. That is an engineering decision, an operational decision, and increasingly a business decision.
Reliability Means Planning for Failure
Any technology operating at enterprise scale will eventually encounter problems. AI is no different. Models may become unavailable. External services may fail. Data sources may be incomplete. Outputs may fall outside acceptable thresholds. Agents may encounter situations they were not designed to handle. A mature AI operating model assumes these things will happen and plans accordingly.
That means establishing incident management practices, escalation paths, fallback mechanisms, human intervention points, auditability, and clear ownership. Teams need to know when an AI system should continue operating, when it should degrade gracefully, when a human should take control, and when the capability should stop altogether. The objective is not to pretend failure can be eliminated, but rather to build systems that can recognize, contain, recover from, and learn from failure.
Governance Does Not End at Deployment
The same principle applies to governance. Security, privacy, compliance, responsible AI, and human oversight cannot simply be reviewed before deployment and then considered complete. Production creates new information about how an AI capability is actually being used. Employees may use it differently than expected. New data may enter the workflow. Agents may gain access to additional tools. Regulatory requirements may evolve. New risks may become visible only after the system has been operating for some time.
Governance therefore needs to remain connected to operations. This does not mean surrounding every AI system with unnecessary bureaucracy. It means making governance proportional to the risk and embedding appropriate controls into the way the capability operates. For low-risk applications, that may be relatively lightweight. For systems making consequential decisions or taking autonomous actions, the requirements will naturally be more significant. The principle remains the same: trust has to be maintained, not simply established.
Keep People in the Loop Where They Matter
Human oversight is often discussed as a general principle of responsible AI. In practice, it needs to be much more intentional. Organizations should understand where human judgment adds value, where it is required because of risk, and where it simply creates unnecessary friction. AI Operations provides the evidence needed to make those decisions. If humans consistently override a particular type of output, that tells us something. If an agent repeatedly escalates the same scenario, that tells us something. If people routinely ignore a recommendation, that tells us something too. Those patterns should become inputs into the improvement process. The goal is not to maximize human intervention or eliminate it. The goal is to design the right relationship between people and intelligent systems for the workflow being performed.
Operations Should Make the Knowledge System Smarter
One of the most valuable outputs of AI Operations is not simply keeping systems running. It is what the enterprise learns from operating them. Every production system generates knowledge. Teams learn which models perform best for particular workloads. They discover effective monitoring patterns, security controls, escalation processes, cost-optimization techniques, architectural approaches, governance practices, and ways of integrating human judgment. That knowledge should not remain trapped inside an individual project team.
Within our Engineering Intelligence Knowledge System, those lessons become reusable playbooks, runbooks, patterns, reference architectures, governance practices, monitoring approaches, and other assets that can improve future engagements. This creates another part of the compounding loop within the Enterprise AI Operating System. Production generates evidence. Evidence creates learning. Learning improves the knowledge system. And, that knowledge makes the next solution easier to design, deploy, operate, and improve.
Going Live Is the Beginning
As enterprises move from AI experimentation toward AI-native operations, Day 2 will become just as important as Day 1. Organizations will need to operate AI with the same seriousness they apply to other mission-critical enterprise capabilities, while recognizing that intelligent systems introduce new requirements around model behavior, data, economics, observability, governance, and human oversight.
The companies that do this well will not simply deploy more AI. They will build an organizational capability for continuously operating, measuring, learning from, and improving it. Because ultimately, the real test of an AI system is not whether it worked during the demonstration., but whether it continues to work when it encounters real people, real workflows, real data, real costs, and real business consequences. Going live is not the end of the AI journey. It is where the enterprise begins learning what the technology can actually do.