Operator's playbooks 6 min read

Five questions to ask before you buy another AI tool

Most AI purchases are evaluated on capability, which is the least useful thing to evaluate them on. These five questions take about twenty minutes and will kill roughly half of what crosses your desk, which is the point.

The pattern is familiar to anyone who has sat through a few of these. A demo is genuinely impressive. Everyone in the room agrees it is impressive. A pilot gets approved on that basis, and a quarter later nobody can say whether it worked, so it quietly renews or quietly does not.

The problem is that impressiveness is not a purchasing criterion. Nearly every current tool demos well, because the underlying models are good and the demo was built to show them at their best. What separates the projects that return money is entirely in the questions asked before the purchase, and those questions are about your business, not the product.

Here are the five we use. Ask them out loud, in a room, and write down the answers.

1. What number will this move, and what is that number today?

Both halves matter, and the second half is where most proposals collapse. Naming an outcome is easy. Stating its current value requires someone to have actually measured it, and frequently nobody has.

If the answer is a category rather than a number, that is a finding. "It will improve productivity" is not an answer. "It will cut the time our team spends building the weekly account report, which is currently about eleven hours across four people," is an answer, and it comes with an implied dollar figure and a way to check later whether it happened.

Write the current number down somewhere permanent before anything is built. This single habit does more for AI credibility inside a company than any other practice we know of, because it means wins can be proven and losses can be admitted quickly.

If nobody can tell you the starting number, the project cannot succeed. It can only feel like it did.

2. Whose daily work changes, and have we asked them?

Every AI tool changes somebody's job. The question is whether that person has been in the conversation or is going to find out by email.

This is not a courtesy point. The people doing the work know exactly which parts of it are wasteful and which parts look wasteful but are load-bearing. They know the exception cases, and AI projects die on exception cases. Skipping them means discovering all of it during rollout, when changing the design is expensive and goodwill has already been spent.

There is a specific test worth applying. Ask what gets removed from someone's workflow when this goes live. If the honest answer is that nothing gets removed and this is added on top, the tool will not be used past month two, no matter how good it is. Work that gets heavier does not stick.

3. What actually happens to our data?

Ask this plainly and get the answers in writing. Where is data processed, is it retained, is it used to train models available to anyone else, who at the vendor can access it, and what happens to it if you leave.

The answers vary widely, and the good vendors answer immediately and specifically. Vagueness here is itself the signal.

Two things are worth knowing. First, for genuinely sensitive material there are private and on-premise options where nothing leaves your environment, and they are more accessible than most mid-market leaders assume. Second, this question is not just about compliance. Teams quietly refuse to use tools they do not trust with customer or financial data, and that refusal usually shows up as unexplained low adoption rather than as an objection anyone voices.

The version of this question people forget

What happens to the outputs? If your team builds prompts, workflows, and configurations inside a vendor's platform for a year, find out now whether that work is portable or whether you have quietly built your operating knowledge into something you cannot take with you.

4. Who owns this after the pilot ends?

The most common answer is a list of people, which means nobody. IT owns the integration. A sponsoring VP owns the relationship. Two enthusiastic people own the actual usage until they get busy or leave.

What is needed is one named person accountable for the number from question one, with enough authority to change how the work is done. That second part is where this usually fails: a coordinator can run a tool, but only someone with standing can tell a team that its process is changing. Assigning ownership to someone without that authority is a way of appearing to have an owner.

The related question is what happens when that person leaves. If the knowledge to run this lives entirely in one head or entirely with an outside vendor, you have bought a dependency rather than a capability. Decide deliberately which of the two you want, because both are legitimate choices and only one of them is usually what people think they are buying.

5. What would make us stop, and when do we decide?

Almost nobody sets kill criteria, which is why so many pilots neither succeed nor end. Agree in advance what result would mean this did not work, and put a date on the decision.

This protects you twice. It stops sunk cost from keeping a weak project alive, and it protects a genuinely promising project from being judged prematurely by someone who wandered into a status meeting. A written criterion and a written date convert a political conversation into an arithmetic one.

A reasonable default for a first use case: a defined improvement against the baseline within ninety days, reviewed on a fixed date, with an explicit decision to scale, adjust, or stop. Say the word stop out loud during the kickoff. It changes how seriously everyone treats the baseline.

A pilot without kill criteria is not a pilot. It is a subscription with a trial period attached to it.

What to do with the answers

Run the five questions on whatever is currently in front of you. If three or more come back vague, the problem is not the tool. It is that the use case was never defined well enough to buy against, and buying anything now will produce the same result as the last time.

The fix in that case is not a better vendor. It is going back and doing the work of picking a specific process, sizing it in dollars, and deciding who owns the outcome. That work is unglamorous and it is the entire difference between AI spend that compounds and AI spend that renews out of habit.

Where do you stand

Score yourself in two minutes.

The AI Readiness Scorecard runs a longer version of this check across strategy, data, workflow, capability, and governance, then names your two biggest gaps. Free, no login.