This is the fifth piece in the series, and it picks up where the last few left off.
We looked at why most companies are collecting demos instead of adopting GenAI. Then what AI-native architecture actually means, whether your architecture is ready, and where your risk posture stands.
Now the moment every one of those articles was building toward:
The pilot worked. Users like it. Someone senior just asked, “So when do we go live?”
Here’s the uncomfortable truth:
A successful AI demo proves that something is possible. Production proves that it is dependable.
Those are two very different claims, and most teams carry the evidence for the first into the decision about the second.
Three stages, three different claims
Demo: “It works.” Curated inputs, controlled conditions, the builders in the room.
Pilot: “Users like it.” Real users, but a forgiving audience at low volume. Nobody actually depends on it yet.
Production: “We can trust, operate and afford it.” Messy inputs, real scale, audit questions, and a failure at 2 a.m. with nobody from the project team online.
Each stage needs different evidence. Thirty years in IT has shown me that the projects that stall aren’t the ones with a weak model. They’re the ones that never noticed the evidence had changed.
So before go-live, I’d want five questions answered.
1. How do we know it’s still working?
Evaluation and observability
A demo is judged by impression. Production needs measurement.
- Evaluate on reality. Build the test set from real, messy, edge-case inputs, not the examples that made the demo look good.
- Define “good” in numbers. Accuracy, groundedness, escalation rate. If you can’t measure it, you can’t manage it.
- Re-run evaluations on every change. Model, prompt, data or retrieval layer. Any of them can shift behavior silently.
- Log enough to trace. Prompt version, retrieved context, response, latency, token cost and user feedback, so a bad output can be traced back to its cause.
If you can’t detect degradation, your users will detect it for you.
2. What happens when it fails?
Reliability, fallback and failure handling
Models time out. Providers have outages. Retrieval returns nothing. Outputs arrive malformed.
None of this is exotic. It’s Tuesday.
In article 2, I said the practical test of AI-native design is what happens when the model is unavailable, wrong, slow or expensive. In production, that test gets run for you, live.
- Design a path for every failure: a simpler model, a cached answer, a rules-based flow, or a clean handoff to a person.
- Set timeouts, retries and circuit breakers deliberately, not by default.
- Decide what “unsure” looks like. “I don’t know, here’s who can help” is a feature, not a defect.
Reliable systems don’t avoid failure. They make failure designed, not discovered.
3. What does it cost at 100x?
Cost control
A 50-user pilot hides the economics. Production exposes them.
- Model cost per transaction, then multiply by realistic volume, not pilot volume.
- Route by difficulty. Smaller models for easy requests, larger ones for the hard cases.
- Cache, cap and alert. Repeated queries, context size, and budget thresholds per feature.
- Tie cost to outcomes. What does one successful result cost, including monitoring, evaluation and human review, and is that acceptable?
If the numbers only work at pilot scale, you don’t have a product yet. You have a demo with a budget problem.
4. What can it touch, and what if it’s manipulated?
Security
I covered this in depth in the last piece, so here’s the production-gate version:
- Least privilege on every tool, API and data source the system can reach
- Prompt injection defenses wherever it reads external content or takes action
- Sensitive data kept out of prompts, logs and outputs
- Model outputs treated as untrusted input before they reach downstream systems
The more autonomy you grant, the more your security model has to assume someone will try to manipulate it.
5. Who owns it, and where do humans step in?
Human oversight and ownership
This is the question that gets left unanswered most often, and the one I’d weigh most heavily.
- One named owner accountable for behavior, cost and risk. A project team that disbands after launch is not an owner.
- Human review based on impact and confidence, not convenience.
- Reviewers who can actually decide. Give them context and time. A rubber stamp on paper isn’t oversight.
- An incident process: who gets paged, who can switch it off, and how lessons flow back into the system.
Without ownership, every other answer on this list decays within months.
The go-live test
Ask your team to explain, in plain language:
- How do we measure quality?
- What happens when it breaks?
- What does it cost at scale?
- What can it access?
- Who is accountable?
If any answer is “we’ll figure it out after launch,” you’re still in pilot.
The real difference
The demo earned you the meeting. Production earns you the trust.
And the gap between them isn’t a better model. It’s better engineering, clearer accountability, and the discipline to treat AI as a system you operate, not a feature you ship.
Which of these five questions did your team struggle with most before going live? I’d genuinely like to hear it. Drop a comment or send me a message.