The fastest way to tell a real AI implementation from hype is specificity: a real one names a workflow, a tool, an all-in cost, and a measurable outcome, while hype stays at the level of a category and a percentage.
The tell is specificity. Real AI implementations have a workflow, a tool, a cost, and a measurable outcome. Hype has a percentage improvement, a transformation narrative, and no invoices. The difference is consistently visible if you know which questions to ask — and consistently invisible if you don’t, which is why the same pitch works on the same audience indefinitely.
There is a specific kind of AI content — a case study, a webinar, a consultant’s pitch deck — that follows a formula so consistent it might be procedurally generated. A business, unnamed or barely named, deployed an AI solution, unspecified or barely specified, and saw a 40% improvement in something measurable-sounding. The proof is a quote from a satisfied executive. The methodology is not discussed, because there isn’t one.
This format persists because it works: vague enough to be unobjectionable, specific enough to sound real, optimistic enough to justify the budget line.
The antidote, it turns out, is five questions.
The five questions
1. “What workflow, specifically, does this automate?”
A real answer names a process. "New client intake emails." "First-pass invoice coding." "Dispatch ticket creation from voicemails." A hype answer names a category: "customer communications," "back-office operations," "the sales process." The category answer tells you they've built a demo, not a system.
2. “What does a failure look like, and how often does it happen?”
Anyone who has shipped an AI implementation knows exactly what the failure mode is — because they've watched it fail. The model misreads a voicemail with background noise. The prompt works on English text and produces nonsense on anything else. The API rate-limits under load and nobody noticed until three days of tickets went missing. If they can't describe a failure, they're describing a controlled environment, which is to say, a demo.
3. “What's the all-in cost: tools, setup, and your own time to maintain it?”
The subscription price is a number. The cost is a different number. A $49/month tool that took 80 hours to implement and requires a half-day of maintenance every time the source system changes is not inexpensive. The all-in figure — tools, setup labor at your time's value, ongoing maintenance — is the only number that lets you calculate a payback period.
4. “Can I talk to a business running this in production — not in beta?”
"In beta" means the vendor is still discovering the failure modes, and you are the instrument of discovery. "In production" means it has run on real data, failed in some number of ways, been fixed, and is now reliable enough that someone is depending on it. These are different claims.
5. “What did it not work for?”
This is the most reliable signal of all. Anyone who has honestly shipped AI knows precisely where it breaks — because they found out the hard way. If they answer immediately and specifically, they've shipped something. If they deflect or say it works for everything, they haven't — or they're not telling you about the part where it didn't.
The patterns worth recognizing
"Our AI handles 90% of cases automatically"
Ask what the 10% is and who handles it. If they don't know, the 90% figure is also invented.
"Most businesses see 3–5× ROI"
Ask for the methodology. "Most businesses" is doing a remarkable amount of work in that sentence.
"It'll take about a week to implement"
Ask what happened in the last three implementations. Week-long timelines become month-long timelines when they encounter your actual systems.
"We use GPT-4 / Claude under the hood"
The model is not the product. The workflow integration is. Plenty of catastrophically bad AI products run on excellent models.
What does a real AI implementation look like?
An AI implementation worth taking seriously has all of the following. Not most. All.
- ✓A named workflow (not a category)
- ✓A specific tool stack with version or plan
- ✓An all-in cost: setup + monthly + maintenance time
- ✓A measurable outcome: hours or dollars
- ✓A named failure mode: "it breaks when..."
- ✓A payback period the operator calculated themselves
If a pitch doesn’t have these, it’s not ready for your money. It might be worth watching — some of the most promising tools are genuinely early — but watching costs nothing and spending costs something.