I've spent 20+ years in enterprise software, and the last several building and leading AI chatbot programs for large product portfolios. This is what that work actually looks like, and what I believe about how AI should be run in production.
I'm a Senior Principal Software Engineer with two decades of enterprise delivery behind me. That means platform architecture, full-stack engineering, and production systems that had to hold up under real scale, not demo conditions. That background is why I ended up leading AI chatbot programs instead of just building them. The hard part was never getting a model to respond. It was getting an entire organization's architecture, content, and operations to support a system that talks to customers on its own.
I've led enterprise chatbot programs from architecture through production rollout across multi-billion-dollar product portfolios, including Alta CoPilot, an enterprise copilot initiative I helped ship by partnering across product, engineering, and content teams. That meant owning RAG strategy, retrieval quality, guardrails, model-provider decisions, telemetry, and the unglamorous work of measurably improving response quality release over release, not just at launch but as the product, the content, and the user base kept changing underneath it.
Alongside that, I do AI documentation readiness, localization readiness, and Gartner submission consulting: structured content models, metadata and taxonomy governance, reusable topic libraries, and analyst-ready evidence pipelines. I don't think of that as a separate skill set from chatbot leadership. It's the same skill set. A chatbot is only as trustworthy as the documentation and content system feeding it, and most organizations get that backwards. They buy a model and hope the content problem sorts itself out. It doesn't.
Here's the belief everything above is built on: the model is not the product. The model is a probabilistic reasoning layer, and it will always behave probabilistically. That's not a bug you patch. It's what the technology is. If your strategy for reducing hallucinations is "wait for a better model," you will wait forever and ship something unreliable in the meantime. The fix isn't a smarter model. It's a system of governance and provenance built around the model that constrains what it's allowed to say and lets you verify why it said it.
The AI is probabilistic. The system around it cannot be.
Most "model failures" I've debugged in production were never model failures. They were retrieval failures, stale content, duplicate topics competing for the same query, or metadata that didn't say what it needed to say wearing a model failure as a costume. Hand the model the wrong chunk, or three slightly different versions of the same answer, and it will confidently produce a wrong one, and it will do it well. Improving the retrieval layer, meaning structured content, clean taxonomy, and one governed source of truth per topic, moves response quality more reliably than swapping models ever does.
Governance gets treated as a compliance tax, something bolted on after an incident. I build it in from day one, as infrastructure: structured content models, metadata and taxonomy standards, reusable topic libraries, ownership and review cycles. That structure is what removes the ambiguity that causes hallucinations in the first place. An AI system without governed content isn't a lean, fast system. It's a system that's guessing, and eventually it guesses wrong in front of a customer.
The question I want any AI-generated answer to be able to answer is: why does this say that? Not "which model generated this text," but which ticket, specification, decision, or approved source is behind the statement, and what happens downstream if that source changes. That chain, source to document to answer, is what turns an AI response from a plausible sounding guess into something an enterprise can actually stand behind. It's also the belief behind Pelcrow, the documentation integrity layer I built to operationalize exactly this: structure, memory, provenance, and governance around AI-authored content, so the system knows not just what it's saying but why.
Every chatbot program I've run has had a measurement layer from the start: response-quality scoring, escalation and deflection rates, and tracking for where the system is drifting or degrading over time. Hallucination rate isn't a number you estimate once at launch. It moves as content changes, as usage patterns shift, and as the model provider pushes updates you didn't ask for. If you're not measuring it continuously, you find out about it from a customer complaint instead of a dashboard.
Guardrails belong in the architecture from the first design review, not added after something embarrassing happens in production. That includes knowing where the system should hand off to a human instead of answering on its own. Governance isn't only about constraining what the AI says. It's also about knowing when it shouldn't be the one answering at all.
None of this is anti-AI. I've built my career on shipping it at scale. It's a belief that AI leadership means owning the system around the model: the content, the governance, the provenance, the measurement. That system, not the model, is what determines whether an enterprise can trust what its AI says.
I help teams get chatbot architecture, governance, and provenance right from the start.
Get in touchI help organizations deliver production AI chatbots, AI-ready and localization-ready documentation, and Gartner-grade submissions across global product lines.