How to evaluate a bespoke AI agent provider

We build AI agents for a living, so we’re hardly neutral. But these are the ten questions we’d want a prospective client to bring to a first meeting with us, and if you only have a minute, they’re all you need. If a provider can’t answer them, that tells you something. If they’ve never been asked, that tells you something too.

  1. What does your agent do when it doesn’t have the information it needs? Tell us about a time it happened.
  2. Which parts of our problem shouldn’t be automated at all?
  3. What result would make you stop and redesign?
  4. Which parts can current technology do well, which are marginal, and which can’t be done yet?
  5. Who writes the prompts, and how will you learn our domain?
  6. How will a non-technical manager see what the agent did last week, and why?
  7. Where does our data live, and is anything trained on it?
  8. Show us a model you’ve swapped out.
  9. If we end the relationship, what do we keep?
  10. What’s the oldest agent you have that still runs every day?

If you’re curious about where the questions come from, read on.

Why the first question matters most

A few weeks ago we gave three copies of the same model the same brief and asked each to draft a proposal. Where the source material existed, the three drafts agreed almost word for word. Where it didn’t (the price, the architecture, who the stakeholders were), they went their separate ways, each with complete confidence. Nobody had told them the price, so they made one up. Three different ones.

We have a standing rule that an agent asks for what it doesn’t have and never invents a number, and this particular chain was breaking it. The test exists to catch exactly that, and it did. But it’s a decent picture of the whole problem with buying AI agents: the demo looks fine, and the interesting question is what happens when the system doesn’t know something, and whether anyone is watching when it does.

So ask for a story, not a slide. The time their agent didn’t have what it needed, what it did, and what they changed afterwards. Ours involves a proposal agent that kept inserting a SWOT analysis where none belonged. It had learnt the shape from old proposals and couldn’t tell which parts of the past belonged to the present. We found it by reading the output. There was no metric that would have caught it.

A provider without a story like this hasn’t shipped much, or isn’t telling you.

Which parts of this shouldn’t be automated?

A consultancy we work with drew a line we’ve since borrowed for everything: automate the mechanical phases, leave alone the phases that rest on one person’s judgement and relationships, and add data only where it makes the human work more valuable. Putting a dashboard behind a conversation that works because of who is having it would damage the conversation.

A provider who can say plainly where AI doesn’t belong in your process is worth more than one with a longer capability list. And when a team gets 30 per cent faster, the reflex is to cut 30 per cent of the team. The more useful question is what you can now do that was too slow or too expensive to attempt before. No vendor can answer that for you, but you’ll know within twenty minutes whether they can think about it with you.

What would prove you wrong?

Every agent is a hypothesis: that a piece of work can be done well enough, reliably enough, by a system rather than a person. A good provider can tell you what would show the hypothesis is false. “Ninety per cent of requests handled unaided, and if it’s under seventy after four weeks we stop and redesign.” A provider who offers “it will drive efficiency across the workflow” instead will report to you in the same language once it’s live.

Ask, too, what current technology can’t do for your problem. The honest answer is a list: this works now, this is marginal, this isn’t achievable yet. Things that are unsolvable this year tend to become routine within twelve months, so the honest answer also has a date on it. A provider who says everything is possible is either new or selling.

Who writes the prompts?

It sounds like a detail. An agent is software, and it’s also writing, and the number of people who are good at both is small. Ask whether the team includes people whose job is the language: the instructions, the tone, the judgement encoded in words. If the engineers handle that, expect agents that work in the demo and grate in production.

Ask how they’ll learn your domain, too. On one engagement we weren’t the experts in the client’s matching problem and said so; we built the logic by trial and error and explained every piece of it before it went anywhere near their product. The client, meanwhile, didn’t fully understand the logic producing the results. The fix was a session in a room, not a feature. There will be a gap like that in your case as well, in both directions, and it’s worth knowing early how the provider intends to close it.

What will we be able to see, and what will we keep?

Once it’s running, can a non-technical manager review last week’s output and understand why the agent did what it did, without an engineer beside them? If visibility is promised for phase two, be careful. Every failure we’ve caught in our own agents, we caught by looking.

Then the ownership questions, which everyone leaves for the contract stage and which matter most three years on. Where does your data live, during and after. Whether anything is trained on it. Whether the models can run inside your own environment if security asks, and whether they’ve actually done that. Whether the agents depend on one model vendor: our own rule is that a vendor-specific feature earns its place only if it does something impossible elsewhere, otherwise it’s a future migration. Ask to see a model they’ve swapped out, not to hear that they could.

And what you keep if you part ways. Code, prompts, integrations, and whatever the agent has learnt about your business in the meantime. A provider who is confident in the work makes it easy to leave.

How often will we see something running?

Bespoke agent work is research as much as engineering, and a fixed-price, fixed-scope proposal for an unsolved problem is either padded or optimistic. Monthly working software is reasonable. Quarterly is a warning. “At the end of the project” means you’re funding a bet you can’t see.

Ask about longevity while you’re at it. The least glamorous agent we run is a three-minute daily podcast for a consumer brand. It has produced an episode every morning for over a year without anyone touching it. Nobody puts that in a keynote, and it’s the best evidence we have.

The answer you’re really after

The best answer to all ten questions is a provider who, somewhere in the first conversation, tells you that part of what you asked for isn’t worth building and suggests something better. Last month one of our agents evaluated three opportunities that another agent had found and rejected two of them. It was the most reassuring thing it did all week.

Activate Intelligence builds AI agents and company memory platforms for financial services, professional services, and corporate communications teams. If you have a proposal on your desk and would like a second opinion, we’re happy to give one. Talk to us.

Research notes, by email

We publish what we learn – a few pieces a month, no noise. Get the next one in your inbox.

Tags:

Leave a Reply

Discover more from Activate Intelligence

Subscribe now to keep reading and get access to the full archive.

Continue reading