A demo is built to succeed. Your operation is built to expose. The distance between the two is where most AI programs quietly die. Here are the questions that tell you which one you are being sold, before you sign.
Every AI demo works. That is what it is for. The inputs are clean, the path is happy, and the hard cases are not in the room. Your operation is the opposite. It runs on the messy inputs, the edge cases, and the days when volume spikes and the data is wrong. The gap between the demo and the operation is where most enterprise AI quietly dies, and independent research keeps putting the share that reaches production in the single digits.
I sit on three sides of this. I evaluate large language models for a frontier AI lab, so I know how these systems fail. I run a team of AI agents in production, so I know what it takes to keep them honest. And I have run the contact centers and insurance operations these tools get pointed at, so I know what the demo leaves out. I have written before about where AI earns its place in an operation. This is the buyer's side of that question: how to tell, before you sign, whether the thing in the demo will survive your floor.
A demo runs on curated inputs. Ask for the ugly ones: the accented caller on a bad line, the half-scanned ACORD form, the policy with three endorsements and a mid-term change. If the vendor will only show you the clean path, the clean path is what works. The exceptions are not a rounding error in an operation. They are where the cost and the risk live.
Every real system fails. The question is whether the failure is caught or shipped. How does the system know when it is wrong. What is the fallback when confidence is low. Who catches the miss, and how fast. The answer that should worry you most is "it doesn't really fail," because a system that cannot recognize its own miss will hand you a confident wrong answer and call it a success. I see that failure mode every week in evaluation work, and it does not announce itself.
A demo shows capability. An operation shows a number. Faster, better, and cheaper are three separate claims, and a good vendor will prove each one on your data, against your quality bar, not on a benchmark that flatters the tool. Ask for cost per resolved contact or per clean transaction, not automation rate. Automation rate is easy to raise and easy to fake. The resolved outcome is the thing you actually buy.
Every real deployment has an exception path, and the people on it are part of what you are buying. A solution has that path designed, staffed, and measured. A demo pretends the exceptions are not there. Ask what happens on handoff, how the person gets the context, and how the queue is staffed for the cases the model declines. The AI and the human coverage behind it are one system, not two line items.
The demo is launch day. The cost and the risk are in month six. Models drift, data pipelines break quietly, and a system that was accurate in the pilot degrades if no one is watching. Who monitors it, who retrains it, who owns the data staying clean, and is that priced into the number you were quoted. A launch is a project. Running it is the job.
The tells are in the answers. A team that has run operations answers in specifics: the edge cases, the fallback, the number, the staffing behind the exception. A team that has only built demos answers in adjectives. Flowery and factless usually means they have not run it, and a vendor who has not run it will learn on your operation, on your budget.
These are the questions a careful buyer asks. They are the same ones the person who owns the result asks, only harder, because they will still be there in month six. Someone who sells the AI and moves on can afford a demo. Someone who closes the deal and then owns the account through implementation, go-live, and the years of operation after cannot, because every gap the demo hid becomes their problem on a Tuesday. If you are bringing AI into an operation that is not built around it yet, that accountability is the whole game. You do not start with the moonshot. You start with one use case where a mistake surfaces fast and the number is measurable, prove it in production, and earn the right to do the next one. Growing an account with AI is a sequence of kept promises, not a platform sale.
There is no perfect answer for every operation; it's a mix of the tool, the people, and the process, designed for the specific work in front of you. However, whether you are buying AI or you are the one who will have to make it work in a client's operation, the difference between a program that ships and one that dies is usually decided before the contract, in the questions someone was willing to ask. You get what you pay for, and you find out what you paid for in production.
If you are weighing an AI pitch, or planning how to bring AI into an operation that is not built around it yet, and you want a second set of eyes from someone who evaluates these models, runs them in production, and has owned the accounts and the SLAs they get pointed at, that is worth at least having the conversation.
Sources: production-rate context from the MIT NANDA "State of AI in Business 2025" report and Gartner's 2025 estimate on agentic AI project cancellations.
Tell us what you are trying to fix, scale, or evaluate. We will give you an honest read on whether we can help.
Contact us