Gartner has a name for something happening across the AI vendor market right now.
They call it agent washing.
It means ordinary chatbots and automation scripts get relabeled as agentic AI. None of the real autonomy the term implies is actually there. Of the thousands of companies marketing agentic AI today, Gartner estimates only a small fraction, somewhere around 130, genuinely offer it.
Gartner predicts that more than 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls.
Most of that failure rate is preventable. It gets decided long before launch, at the evaluation stage most companies rush through.
Here is a real framework for doing that stage properly, in the order it actually needs to happen.
One distinction is worth making before any of this.
This framework is built for evaluating a product vendor, a company selling you a pre-built AI platform or tool you will use as-is.
Hiring a development partner to build a custom AI feature for your own business is a related but different decision, since ownership, data rights, and lock-in work differently when the thing being built is yours from the start.
We fall into that second category ourselves, and we think the same honesty this framework asks of a product vendor should apply just as much to a build partner. More on that at the end.
Vendor selection and contract signing are not one decision. They are two decisions, and they fail in different ways. Selection is about choosing which company earns a real trial. That choice rests on capability, fit, and reputation.
Signing is a separate, later step. It focuses entirely on the legal terms, the data rights, the pricing structure, and the exit conditions written into the agreement.
A weak selection process wastes an evaluation cycle. A weak contract locks a business into bad terms for years. Treating both as the same conversation is how a company ends up with a great vendor and a terrible agreement.
Before comparing features, ask a more basic question. Does this system actually take independent action? Does it use tools, make its own decisions within set limits, and adjust based on results? Or does it just follow a fixed script with an AI-written response layered on top?
A genuinely agentic system can be tested directly. Ask a vendor to show it handling a task that needs a real decision partway through. Not a smooth, pre-scripted demo path. If the vendor cannot show that without weeks of custom setup, the honest answer is usually that the product is not there yet.
A vendor demo is a controlled performance. The vendor picks the data. The vendor controls the conditions. The vendor shows the best possible version of the system. None of that tells you how it performs on your actual work.
Insist on a proof of concept using real data from your own business. Agree on the success measure before the trial starts. Do not adjust it afterward once results come in. A vendor unwilling to agree on that measure in advance is telling you something worth hearing.
Reference customers are one of the more reliable signals of vendor quality. More reliable than funding announcements or analyst rankings. Ask for direct conversations with at least three customers similar to your own business, in industry, size, and use case. Ask them what the vendor cannot fully answer for you: what broke, what took longer than expected, what they wished they had asked before signing.
| Question | Why it is Separate Issue |
| Is our data stored securely? | A security certification answers this question, and only this one |
| Can the vendor train on our data or outputs? | A completely separate contractual right, often granted by default unless specifically opted out |
These get treated as one question constantly, and they are not. A vendor can store data securely, meeting every relevant certification, while still keeping the right to use that same data, or what the system generates from it, to improve their own models. Confirm in writing that your data and outputs stay yours. Not licensed back to the vendor by default.
Ask exactly how data export works if the relationship ends. Ask what format it comes in, and how long it takes. Ask what happens if the vendor changes the underlying model without warning, a real and increasingly common scenario as providers update their own systems. A contract silent on either point is a contract written entirely in the vendor’s favor.
This is the step most companies skip. It is also the one that decides whether everything above actually holds up. We argued this at length in every failed tech project has the same root cause, and it is never the technology.
The tool is rarely what fails. The absence of a specific, named, accountable owner is what fails. That decision needs to be made before the ink dries. Not scrambled together after the system is already live.
Start with a written checklist sent to six to eight vendors. Run live evaluation sessions with three of them. Take no more than two into a real, paid proof of concept. Spreading a full trial across a longer list makes it hard to test any single vendor properly. That defeats the entire point of testing in the first place.
We said earlier that we fall into the build-partner category ourselves, not the product-vendor category this framework is built around.
The honest version of that is simple: the same questions apply to us too, just answered differently, since a custom build means the code and the data are yours from day one, not licensed back to us.
Ask us these seven questions directly if you are weighing us as a partner, or bring us a vendor decision you are trying to evaluate. Either conversation is one we are glad to have.
AI & Enterprise Strategy
Checked the weather this morning? Paid for something online? Used a map app to find directions? Each of those small moments quietly used something called an API. You never saw it happen. That is exactly the point. This piece explains what an API actually is, in plain words, with no coding knowledge needed. What Does […]
AI & Enterprise Strategy, AI Models
You have almost certainly used one. ChatGPT is one. Gemini is one. Claude is one. Every time you type a question into an AI chat tool, one of these systems is doing the work behind the scenes. The name sounds technical. The real idea is simple once someone walks you through it, step by step, […]
AI & Enterprise Strategy
You ask an AI chatbot a question. It gives you a clear, confident answer. It sounds right. But it is not right. The fact is wrong. The source does not exist. The quote was never said. This happens more than most people think. It has a name in the AI world. It is called hallucination. […]