Questions to ask an AI agency before hiring them
12 questions that separate real AI agencies from ChatGPpt wrapper shops. Know what to ask before you sign, and how Klevere answers each one.
Klevere AI Team
AI Strategy
You are about to spend five or six figures on an AI implementation. The agency pitching you has a glossy deck, a friendly founder, and three case studies that all look suspiciously similar. Your CFO wants proof of ROI. Your CTO wants to know what happens when the agency leaves. Your operations director wants to know if this thing will actually integrate with Salesforce, or if it is another science project that dies in a sandbox. You need a list of questions to ask an AI agency that cuts through the pitch and gets to what matters.
This is that list. Twelve questions that separate agencies building real AI agents from the ones reselling a ChatGPT wrapper with a Zapier bolt-on. Each one comes with the honest answer Klevere gives, the red flags to watch for, and why the question matters in the first place. If an agency stumbles on more than two of these, you are likely talking to the wrong shop.
1. What AI projects have you delivered in my industry, and can I speak to those clients?
**Why this question matters.** Industry context changes everything in AI. A recruitment agent that works brilliantly for a staffing agency will fail in law if it cannot handle privilege, conflicts checks, and matter codes. An ecommerce recommendation engine tuned for fashion will produce nonsense results in industrial components. Generic AI demos look impressive until you try to run them against your actual data model, your compliance regime, and your users' daily workflow.
Ask for named case studies in your sector. Then ask to speak to at least one of those clients directly, not through a reference call the agency has stage-managed. You want to hear what went wrong, how long integration really took, and whether the agency stuck around when the first deployment hit a wall.
**Klevere's answer.** We have delivered projects across 12 industries, with public case studies in recruitment, marketing, sales, and data intelligence. Our /case-studies/recruitment-agent work for KlearSkill analysed over 1 million candidates with 95 per cent match accuracy. We have three additional case studies under NDA in furniture and design, recruitment outreach, and VC deal flow. If you are in one of those sectors, we will connect you with a reference client after you sign an NDA. If we have not worked in your industry, we will tell you that in the first conversation and explain how we would approach the domain learning required.
**Red flags.** The agency only has case studies in unrelated industries and insists the work is transferable. They refuse to provide references, citing confidentiality on every project. They offer a 'sample client' to call but will not let you choose which one. They have no NDAs in place and happily name drop clients without permission.
2. Who actually builds the agents, and will I meet them before we start?
**Why this question matters.** AI agent development is not a commoditised service you can offshore to a junior team following a playbook. It requires people who understand your business logic, can write reliable code that handles edge cases, and know when to push back on a requirement that will break in production. If the people in the sales meeting are not the people doing the build, you need to meet the build team before you sign anything.
Some agencies operate a partner model where they sell the work, then subcontract delivery to freelancers or offshore teams they have never worked with before. Others have a rotating bench of contractors. Both models can work if the agency is honest about it and has strong technical leadership holding the delivery quality line. The failure mode is when you discover halfway through the project that the person who sold you on their deep expertise is not actually touching your code.
**Klevere's answer.** Every Klevere project is led by a named AI engineer and a product lead, both of whom you meet during the scoping process before any contract is signed. We do not operate a subcontractor model. The people who build your agent are Klevere employees, and they stay on the project from discovery through to deployment and handover. We will introduce you to your project team during the proposal stage, and you can request CVs if that is part of your vendor diligence process.
**Red flags.** The agency will not name the team until after contract signature. They describe their delivery model as 'flexible' or 'scalable' without explaining who is actually writing code. The salesperson is evasive when you ask if they personally do any of the technical work. The agency talks about 'our network of specialists' but cannot tell you if those specialists are employees, contractors, or partners.
3. What happens when the agent breaks or hallucinates in production?
**Why this question matters.** Every AI agent will produce a wrong answer at some point. The question is whether the agency has designed the system to catch that failure before it reaches a customer, a regulator, or a revenue-critical process. Hallucination is not a problem you solve by picking a better model. It is a problem you solve with guardrails, validation layers, human-in-the-loop workflows, and monitoring that alerts you when output confidence drops below a threshold.
What you want to hear is a specific technical answer about how the agent is instrumented, what monitoring is in place, and how quickly the agency can patch a failure mode once you report it. What you do not want to hear is a vague reassurance that the model is very accurate and problems are rare.
**Klevere's answer.** Every agent we build includes confidence scoring on outputs, validation rules that catch malformed or out-of-bounds responses, and structured logging that lets us trace exactly what the agent did and why. For high-stakes workflows like contract review or financial reconciliation, we design mandatory human-in-the-loop steps where a person approves the agent's recommendation before it executes. If something breaks in production, our SLA is a four-hour response time for severity-one issues, and we provide a root cause analysis within 24 hours. We also run regular adversarial testing against your agent to find edge cases before your users do.
**Red flags.** The agency says their models do not hallucinate because they use the latest version. They suggest you just need to write better prompts. They have no monitoring in place and rely on users to report problems. They cannot explain what happens if the agent gives a wrong answer that costs you money or exposes you to liability. They do not offer an SLA or incident response process.
4. How do you handle data sovereignty, compliance, and security?
**Why this question matters.** If you operate in healthcare, financial services, legal, or any regulated industry, your data cannot leave certain jurisdictions, cannot be used to train third-party models, and must be encrypted both in transit and at rest. If the agency is sending your data to OpenAI's public API without a Business Associate Agreement, you are in breach of HIPAA the moment you go live. If they are storing everything in a US region and you have UK or EU customers, you may be in breach of GDPR.
This is not a theoretical concern. Regulators have started issuing fines for AI systems that violate data residency and privacy rules, and your agency's ignorance is not a defence. You need to know where your data is going, who can see it, and what agreements are in place to prevent it being used for purposes you have not authorised.
**Klevere's answer.** We are SOC 2 Type II and ISO 27001 certified, with HIPAA, GDPR, and CCPA compliance as standard. We offer regional data residency in the UK, EU, US, and Australia, and we can deploy agents that never send data outside your own infrastructure if that is a requirement. All client data is encrypted in transit and at rest, and we use dedicated instances and API keys that contractually prohibit model training on your data. If you need a BAA, DPA, or other compliance agreements, we provide those as part of onboarding. You can see our compliance documentation during the proposal process, and we will walk through any specific regulatory requirements your legal team flags.
**Red flags.** The agency has no compliance certifications and says they can probably sort something out if you need it. They cannot tell you which region your data will be processed in. They use the public ChatGPT API or other consumer-tier services that do not offer enterprise data protection. They tell you GDPR is not a problem because the model does not store data, which shows a fundamental misunderstanding of how privacy regulation works.
5. Can I see the actual code, or do I just get API access?
**Why this question matters.** Some agencies treat AI agents as a black box service. You pay a monthly fee, you get API access, and if you ever want to leave or bring the work in-house, you are starting from scratch. Other agencies build the agent as a work-for-hire, hand over the codebase, and train your team to maintain it. Both models are legitimate, but you need to know which one you are buying before you sign the contract.
If you are paying for custom development, you should own the code. If the agency insists on retaining IP, they need to justify why, and you need to understand what that means for your exit options, your ability to integrate with other systems, and your leverage if the relationship goes sour.
**Klevere's answer.** Every custom agent we build is delivered as a work-for-hire. You own the code, the training data, the prompts, and the configuration. We hand over a fully documented repository, along with architecture diagrams, API documentation, and runbooks that explain how to operate and extend the system. If you want to bring the work in-house or switch to another agency after we have finished, you can do that without friction. We also offer ongoing support and development retainers, but those are optional, not a lock-in mechanism.
For clients using our /ai-os product, the situation is different. The AI OS is a managed platform, and you are licensing access rather than buying the underlying codebase. We are clear about that distinction up front, and we provide export tools so your data and workflows are portable if you ever decide to leave.
**Red flags.** The agency will not discuss code ownership until after you have signed. They claim the IP must stay with them to protect their proprietary methods, but cannot explain what those methods are. They offer only API access and become evasive when you ask about migration or exit. They bundle the agent into a long-term SaaS contract with no early termination clause.
6. What is your process for scoping and pricing a project?
**Why this question matters.** Fixed-price AI projects are a gamble unless the agency has done something nearly identical before. The discovery process almost always surfaces requirements, edge cases, or integration challenges that were not visible in the initial brief. Honest agencies price that uncertainty into the engagement model, either as a phased contract with a discovery stage before the main build, or as time-and-materials with a cap. Dishonest agencies lowball the quote to win the work, then hit you with change requests the moment you ask for something that was obviously in scope from day one.
What you want is a proposal process that includes enough discovery to give you a realistic price and timeline, a clear statement of what is in and out of scope, and a mechanism for handling changes that does not turn into commercial warfare every time you spot a requirement you forgot to mention.
**Klevere's answer.** We do not quote a price until we have run a free 30-minute AI audit with you, reviewed your current workflows, and understood your data landscape and compliance constraints. That audit is available at /solutions/ai-audit and it is the input to our scoping process. After that, we provide a proposal with a fixed price for a defined scope, along with a timeline and a list of assumptions. If scope changes during the project, we handle minor adjustments within the existing budget and flag anything material for a change control conversation. We do not play games with change requests. If we missed something in discovery that should have been obvious, we absorb that. If you introduce a new requirement that changes the architecture, we will re-scope and re-price that work transparently.
We do not publish day rates, build fees, or pricing tiers, because every engagement is scoped and priced individually based on complexity, data volume, integration requirements, and the level of support you need post-launch.
**Red flags.** The agency quotes a price in the first meeting without asking detailed questions about your data, systems, or processes. They insist on fixed-price for a vague scope and refuse to break down the estimate. They have a standard rate card that does not flex for project complexity. They nickel-and-dime you on change requests that were clearly implied by the original brief.
7. How much of this is custom-built versus off-the-shelf components?
**Why this question matters.** There is no point paying for bespoke development if an off-the-shelf tool solves 80 per cent of your problem. Equally, there is no point buying a SaaS product and paying an agency to customise it if the vendor's API is too restrictive to deliver what you actually need. Good AI agencies are honest about when to build, when to buy, and when to integrate.
The failure mode is agencies that have a hammer and treat every problem as a nail. If they only do custom LangChain development, every solution looks like a custom agent. If they only resell a specific platform, every solution looks like that platform plus some light configuration. You want an agency that picks the right tool for the job, even if that tool is not the one that makes them the most margin.
**Klevere's answer.** We use a stack that includes OpenAI, Anthropic, Google Gemini, LangChain, Pinecone, Weaviate, Salesforce, HubSpot, Slack, Microsoft 365, AWS, and Snowflake, depending on what your use case requires. If an off-the-shelf integration solves your problem, we will tell you that and save you the cost of custom development. If you need a fully custom agent, we will build that and explain why the off-the-shelf options do not fit. Most projects land somewhere in the middle, using existing platforms for data storage and orchestration, with custom agent logic and prompt engineering tailored to your workflows. We are not religious about build versus buy. We are religious about delivering something that works.
**Red flags.** The agency insists on custom-building everything, even when there are proven SaaS tools that do the job. They have a commercial partnership with a platform vendor and push that product regardless of fit. They cannot articulate why they chose one model or framework over another. They describe their offering as a fully proprietary platform but cannot explain what makes it proprietary beyond a custom UI on top of standard components.
8. What does success look like for this project, and how do you measure it?
**Why this question matters.** If you cannot measure whether the AI agent is working, you cannot manage it, optimise it, or justify the cost when budget review season arrives. Some use cases have obvious metrics, like a sales agent that books meetings or a support agent that resolves tickets. Others are fuzzier, like a strategy agent that summarises reports or a recruitment agent that screens CVs. In both cases, you need to agree up front what good looks like, and you need instrumentation that tracks it.
Agencies that care about outcomes will ask you what metric you are trying to move, and they will design the agent to optimise for that metric. Agencies that care about deployment will hand over the system, declare victory, and leave you to figure out whether it was worth the money.
**Klevere's answer.** We define success criteria during the scoping process, and we bake measurement into every agent we build. For our /case-studies/recruitment-agent work with KlearSkill, success was 95 per cent match accuracy and the ability to process 1 million candidate profiles. For our autonomous sales agent case study, success was 500 leads contacted and an 85 per cent response rate. We track those metrics in dashboards that update in real time, and we review them with you weekly during the first month post-launch, then monthly after that. If the agent is not hitting the agreed targets, we treat that as a delivery failure and work with you to fix it at no additional cost.
**Red flags.** The agency cannot tell you how they would measure success for your use case. They talk about efficiency and productivity but have no baseline to compare against. They deploy the agent and disappear without setting up any monitoring or reporting. They define success as 'the agent is live' rather than 'the agent is delivering the business outcome you paid for'.
9. What ongoing support and maintenance do you provide, and what does that cost?
**Why this question matters.** AI models degrade over time as the underlying data distribution shifts. Integrations break when third-party APIs change. Compliance requirements evolve and you need to update how the agent handles personal data. If the agency builds the system, hands it over, and walks away, you will need in-house expertise to keep it running, and most SMBs do not have that.
You need to know what post-launch support looks like, whether it is included in the project price or sold separately, and how quickly the agency responds when something breaks. You also need to know what happens if the agency is acquired, pivots to a different market, or just stops answering emails.
**Klevere's answer.** We offer three support tiers: basic monitoring and incident response, which is included in the first three months post-launch; proactive optimisation and monthly performance reviews, sold as an optional retainer; and full managed service where we own uptime, model retraining, and integration updates, also sold as a retainer. Pricing depends on the complexity of the agent and the SLA you need, and we scope that during the proposal conversation. If you want to run the agent in-house after the initial support period, we provide full documentation, runbooks, and a handover session with your technical team. We have a 98 per cent client retention rate, and most of our clients stay on a support retainer because it is cheaper than hiring the expertise internally.
**Red flags.** The agency includes no post-launch support in the project and expects you to figure it out. They bundle support into a long-term contract at a price that is not justified by the level of service. They have no SLA and no incident response process. They cannot tell you what happens if a model is deprecated or an API they depend on changes.
10. How do you handle model updates and changes in the AI landscape?
**Why this question matters.** The AI stack changes every quarter. New models are released, old models are deprecated, API pricing changes, and new capabilities appear that make your current architecture look outdated. If your agent is hard-coded to GPT-4 and OpenAI sunsets that model, you need a plan for migration. If a new model halves your inference costs, you want the option to switch without rebuilding everything.
Good agencies design for model portability. They abstract the model layer so you can swap providers without rewriting the agent logic. They monitor the AI landscape and proactively recommend upgrades when something better becomes available. Bad agencies lock you into a specific model or provider because that is what they know, and charge you for a full rebuild when that choice becomes obsolete.
**Klevere's answer.** We design every agent with a model abstraction layer that lets us swap between OpenAI, Anthropic, Google Gemini, or open-source models without changing the core logic. When a new model is released, we test it against your use case and let you know if it delivers better performance, lower cost, or new capabilities worth adopting. We do not charge for model upgrades if the interface is compatible. If a major architectural change is required, we scope that separately and explain why it is needed. We monitor provider roadmaps and deprecation schedules, and we give you at least 90 days' notice if something in your stack is being sunset.
**Red flags.** The agency builds everything around a single model and cannot explain how they would handle a provider change. They have never upgraded a client to a new model version. They treat model selection as a one-time decision and have no process for revisiting it. They charge full rebuild costs for what should be a configuration change.
11. Can you show me a failure and how you fixed it?
**Why this question matters.** Every AI project hits problems. The model does not generalise the way you expected. The integration takes three times longer than planned because the legacy API is documented incorrectly. The agent works in testing but falls over when you throw production data at it. Agencies that have never had a project go sideways are either lying or have not done enough work to encounter real complexity.
What separates good agencies from bad ones is not whether they have failures. It is whether they can talk about those failures honestly, explain what they learned, and show you how they changed their process to prevent the same mistake twice. If an agency cannot give you an example of something that went wrong and how they recovered, they are either inexperienced or unwilling to be honest with you.
**Klevere's answer.** In one of our early recruitment projects, we built a CV screening agent that worked brilliantly in testing but produced biased results when deployed against a real candidate pool, because the training data over-represented certain universities and prior employers. We caught it during UAT, paused the rollout, retrained the model with a balanced dataset, and added bias detection to our standard QA process for any agent that touches hiring decisions. We now run fairness audits on every recruitment and HR agent we build, and we document those audits in the handover pack. That failure made us better, and we are not shy about talking through it with prospective clients.
**Red flags.** The agency claims they have never had a project miss a deadline or fail to meet a requirement. They become defensive or evasive when you ask about challenges. They blame failures on client scope creep or unrealistic expectations without acknowledging their own contribution. They have no process improvements or lessons learned from past work.
12. Why should I hire you instead of building this in-house?
**Why this question matters.** This is the question every agency dreads, because the honest answer is sometimes 'you should not hire us, you should build this yourself'. If you already have strong AI talent in-house, a clear technical roadmap, and the capacity to take on another project, paying an agency is probably a waste of money. But if you are in that position, you would not be reading this checklist.
Most SMBs do not have AI specialists on staff, and hiring one takes months and costs six figures before you have written a line of code. Most internal IT teams are underwater with BAU work and cannot take on a multi-month AI build without dropping something else. Most organisations do not have the pattern recognition that comes from doing ten or twenty AI implementations and knowing which architectures work and which ones look good in a demo but collapse under production load. That is where an agency adds value, but only if they are honest about when they do not.
**Klevere's answer.** You should hire Klevere if you need AI capability faster than you can hire for it, if you want to validate a use case before committing to a full-time hire, or if you need expertise across multiple domains like compliance, integration, and agent orchestration that would take three or four hires to cover in-house. You should not hire us if you already have a senior AI engineer with delivery experience, if your use case is simple enough to solve with off-the-shelf tools and a few days of configuration, or if your primary goal is to build internal capability rather than deliver a working system quickly. We say no to projects where we do not think we are the right fit, because we would rather lose a sale than take on work that will not deliver the outcome you need. See our /solutions/ai-strategy page for how we think about that trade-off, or book a free audit at /solutions/ai-audit and we will tell you honestly whether an agency engagement makes sense for your situation.
**Red flags.** The agency insists you need them regardless of your internal capability. They dismiss the idea of in-house development as naive or risky without explaining why. They cannot articulate what specific value they bring beyond 'we have done this before'. They sell you a six-month engagement when a two-week pilot would prove the concept.
How Klevere approaches AI agency selection from the other side
We answer these questions to ask an AI agency every week, because every prospective client should be asking them. We built our service model around the assumption that you are doing proper diligence, that you are comparing us to other agencies, and that you will walk away if we cannot back up our claims with proof points, case studies, and straight answers.
That assumption has shaped how we price, scope, and deliver work. We start every engagement with a free 30-minute AI audit because we need to understand your environment before we can tell you whether we are the right fit. We do not pitch a solution in the first meeting. We ask questions, review your systems, and map your requirements to our delivery model. If we are not confident we can deliver what you need, we say so. If you would be better served by a different type of agency, or by building in-house, we will tell you that too.
We have deployed 500+ AI agents across 50+ projects in 12 industries, and we are SOC 2 Type II, ISO 27001, HIPAA, GDPR, and CCPA compliant. Our stack includes OpenAI, Anthropic, Google Gemini, LangChain, Pinecone, Weaviate, Salesforce, HubSpot, Slack, Microsoft 365, AWS, and Snowflake. We have public case studies at /case-studies/recruitment-agent and /case-studies/autonomous-sales-agent, and we can connect you with reference clients in most industries. We also have our /ai-os product for clients who want a managed platform rather than custom development, and a full suite of services at /solutions including custom AI agent development, strategy, automation, and consulting.
If you want to see how we answer these questions in detail, book a free AI audit at /solutions/ai-audit. If we are not the right agency for you, we will tell you that in the first conversation. If we are, you will walk away with a clear scope, a realistic price, and a delivery plan that does not rely on you taking anything on faith.
The AI agency market is full of shops that will tell you what you want to hear, deploy something that looks impressive in a demo, and leave you with a system that does not integrate, does not scale, and does not survive contact with your actual users. The questions to ask an AI agency in this guide are designed to filter those shops out before you waste time and money finding out the hard way. Use them.