What an AI Agent Company Should Show You Before You Sign

A 7-point checklist for hiring an AI agent development company: evaluation sets, cost model, failure behaviour, observability, security, IP and red flags.

By Zoraiz Ejaz, Co-founder, Techparser · · 7 min read

What an AI Agent Company Should Show You Before You Sign

Before signing with an AI agent development company, ask to see seven things: an evaluation set with scores from a past project, a per-request cost model at your volume, a defined failure behaviour (abstain, escalate, cite or gate), a model-agnostic architecture, an observability plan that traces every model call, a written data-handling and security policy, and a contract that assigns prompts, data and code to you.

An AI agent is software that takes a goal, calls tools and models in a loop, and acts on the result. That last part is what makes buying one different from buying an app. A bug in an app shows a wrong screen. A bug in an agent sends the wrong email, approves the wrong refund or cites a document that does not exist. The checklist below is what we ask ourselves before we quote, and what you should ask any vendor before you sign.

The seven things to ask for

1. An evaluation set, with numbers

An evaluation set is a collection of real inputs with known-good outputs, scored automatically or by rubric. It is how anyone knows whether an agent is good enough to launch. Ask the vendor to show one from a previous project, with the score they achieved and how it changed across iterations. If they show you demo prompts instead, they made launch decisions on vibes. Then ask how they will build yours: who supplies the examples, how many, how edge cases are found, and how the set grows after launch from real failures.

2. A cost model per request

Conventional software has near-zero marginal cost per request. Agents do not. Every step calls a model, and multi-step agents can make ten or more calls per task. Ask for an estimate of tokens per task, calls per task and the resulting cost at your expected volume, plus what happens at 10x. A serious vendor will also tell you which steps can use a smaller, cheaper model and which need the strongest one. If the answer is a single blended monthly number with no breakdown, it was not modelled.

3. Defined failure behaviour

The agent will be wrong. The question is what happens next. There are four sound answers, and a production agent usually uses several of them in different places.

  • Abstain: the agent says it cannot complete the task rather than inventing a result.
  • Escalate: low-confidence cases are routed to a human with the context attached.
  • Cite: every factual claim links to a source the user can open and check.
  • Gate: irreversible actions such as payments, deletions or outbound messages require explicit approval.

Ask the vendor to walk through a specific wrong answer in their proposed design and name which of these fires. An agent with no defined behaviour for being wrong is not finished.

4. Model-agnostic architecture

Model quality and pricing shift every few months. The agent should call models through an abstraction so that switching from one provider to another, or to a self-hosted open model for data-residency reasons, is a configuration change and a re-run of the evaluation set. Ask which providers the architecture supports today and how a switch is tested. Our SlimAI food-recognition pipeline runs on Gemini through Firebase; our Spyra Formatter document tool runs on Claude and OpenAI behind the same interface. That is the shape to look for.

5. Observability

You should be able to open any task the agent ran and see every prompt, every tool call, every model response, the latency and the cost. Without that trace, debugging is guesswork and cost overruns are invisible until the invoice. Ask which tracing tool they use, whether traces are retained and for how long, and whether you will have access to the dashboard. Ask how failures are alerted. Ask how a bad trace becomes a new evaluation case.

6. Security and data handling

Agents read your data and act with your credentials. Ask for a written answer to each of these: which model providers see your data and under what terms, whether provider training on your data is disabled, where logs are stored and for how long, how secrets and tool credentials are scoped, how prompt injection through documents or web content is mitigated, and what the agent is technically prevented from doing regardless of what it is asked. If you handle health, financial or children's data, ask how the design meets the relevant regulation before any build starts.

7. IP ownership

The prompts, the evaluation set, any fine-tuned weights, the code and the traces are the product. The contract should assign all of them to you on payment. Watch for agencies that keep the prompts or the orchestration layer as their platform and license it to you. That is a legitimate business model, but it is a subscription, not a build, and you should price it as one.

The checklist as a table

AI agent vendor checklist: what to ask for, what a good answer looks like, and the warning sign
Ask forGood answerWarning sign
Evaluation setScored set from a past project; plan to build yours from real data; grows from production failuresDemo prompts and screenshots; 'we test it thoroughly'
Cost modelTokens and calls per task, cost at your volume and at 10x, cheaper models for simple stepsOne blended monthly number with no breakdown
Failure behaviourAbstain, escalate, cite and gate mapped to specific steps'The model is very accurate'
ArchitectureProvider abstraction; switch tested by re-running evalsSingle vendor SDK called directly from business logic
ObservabilityPer-task traces with prompts, tool calls, latency and cost; alerts; you get dashboard accessServer logs only; no cost per task visible
Security and dataWritten policy covering providers, retention, secrets, injection and hard limits'We use enterprise APIs' with nothing in writing
IP ownershipPrompts, evals, weights, code and traces assigned to you on paymentVendor retains the orchestration layer as their platform

Red flags

  • A fixed price quoted before anyone has seen your data or your tools.
  • Accuracy claims with no evaluation set behind them, or a single accuracy number with no denominator.
  • No human in the loop anywhere in a design that sends messages, moves money or deletes records.
  • A timeline under two weeks for anything that integrates with your systems.
  • No mention of cost per request in the proposal.
  • A demo that only works on the vendor's example documents.
  • Reluctance to put data handling in writing, or a claim that no data leaves your environment when a hosted model is in the design.
  • A contract in which the vendor owns the prompts or the agent framework.

What good delivery looks like

Public 2026 guides from agencies such as GroupBWT put a bounded proof of concept at 2 to 4 weeks, a pilot on a real workflow at 4 to 8 weeks, and a production deployment at 8 to 16 weeks. Cost guides from Neoteric, Biz4Group and others put a proof of concept at roughly $10,000 to $50,000, an MVP with integrations at $20,000 to $100,000, and production agents with memory, permissions and monitoring above $100,000. Treat these as ranges from vendors who sell the work. The right vendor will give you a narrower number for your case after seeing your data, and will tie the first milestone to an evaluation score, not to a demo.

Our AI automation agent service follows this checklist because we wrote it for ourselves first. If a vendor cannot answer these seven questions, keep looking, whether or not that vendor is us.

Frequently asked questions

What does an AI automation agency do?
An AI automation agency designs, builds and operates software agents that complete tasks using language models and tools: triaging support tickets, extracting data from documents, drafting replies, updating records or running multi-step workflows. Good agencies also build the evaluation set, cost model, observability and failure handling around the agent, and integrate it with your existing systems and approvals.
How much does an AI agent cost to build?
Published 2026 guides put a proof of concept at roughly $10,000 to $50,000, an MVP with real integrations at $20,000 to $100,000, and production agents with memory, permissions and monitoring above $100,000. Ongoing inference cost is separate and scales with usage. Techparser quotes per project after reviewing your data and tools rather than publishing a rate.
How long does AI agent development take?
A bounded proof of concept typically takes 2 to 4 weeks, a pilot on one real workflow 4 to 8 weeks, and a production deployment with monitoring and approvals 8 to 16 weeks, according to 2026 vendor guides. The variable is integration: each system the agent reads from or writes to adds authentication, permissions and failure handling that must be designed and tested.
What is an evaluation set for an AI agent?
An evaluation set is a collection of real inputs paired with known-good outputs, used to score the agent automatically or by rubric. It turns 'is it good enough?' into a number that can be tracked across iterations and model changes. It should be built from your data before the UI, and grow after launch as production failures are added to it.
Should an AI agent be allowed to act without approval?
Only for reversible, low-risk actions. Anything irreversible, such as sending money, deleting records or emailing customers, should be gated behind explicit human approval at least until the evaluation set shows sustained accuracy in production. Well-designed agents combine gating for high-risk steps with abstaining and escalating for low-confidence cases.

About the author

Zoraiz Ejaz

Co-founder, Techparser

Zoraiz Ejaz is a co-founder of Techparser and leads its engineering and product practice. He has spent close to a decade designing, building and scaling mobile, web and AI products for startups and enterprise teams across health, fintech, payments, social and education, from first architecture and release pipelines through to launch and years of production support. He writes about how to scope, cost and ship software that lasts.

Related case studies

Related services

Scoped estimate in 48 hours

Tell us what you are building. We reply with scope, timeline and a fixed budget within two business days.

Get a scoped estimate




Keep reading