Posted On: August 14, 2026

AI agent cost is one of the most confusing numbers in technology right now. Ask three vendors for a quote and you will get three different answers, because cost depends on the workflow around the model, not just the model itself: integrations, data preparation, testing, monitoring, human review and which billing model a provider has chosen to use. This guide sets out what build cost, running costs, token and API spend, billing models and realistic returns actually look like for UK businesses in 2026, and why the most useful comparison is cost per successful outcome rather than headline price alone.
Sterling conversions use an illustrative rate of $1.27 to £1, checked in August 2026. They are rounded for comparison, are not live exchange-rate quotations, and should be treated as indicative; always check a vendor's current pricing page before budgeting.
The cost of an AI agent depends less on the language model itself than on the workflow around it: integrations, data preparation, actions, testing, security, monitoring and human review. A ready-made agent may be bought as a subscription or a per-resolution service, while a bespoke, multi-system agent can require a substantial implementation budget running into six figures. Compare vendors using total cost of ownership and cost per successful outcome, not build price or token cost alone.
What determines cost: integrations, data preparation, actions, volume, oversight, service level and frequency of change
Main purchasing routes: ready-made SaaS, configured platform, low-code workflow, bespoke agent, enterprise orchestration or internal development
Biggest budgeting mistake: comparing build prices without including human review, maintenance and usage
Best comparison metric: first-year total cost per verified successful outcome
Best starting point: one narrow workflow with a measured baseline and defined success criteria
Providers price their systems across four very different models: per seat, per token, per task and outcome-based pricing. Each one allocates risk differently between the vendor and the customer, which is exactly why two quotes can look wildly different for what appears to be the same product. A per-seat tool charges a flat monthly fee regardless of usage, while an outcome-based vendor charges only when the system actually resolves something, and those two structures produce very different bills for identical volume.
Part of the confusion comes from treating "AI agent cost" as a single figure. It is really a combination of implementation spend, ongoing model and tool usage, hosting where applicable, and often-overlooked maintenance costs that only surface after launch. A business comparing a £16,000 quote to a £95,000 quote is rarely comparing like for like: one is often a ready-made tool handling a single job, and the other is a connected system spanning several internal platforms and a database, built and supported to a different specification entirely.
"AI agent" is not one purchasing category, and the price bands later in this guide only make sense once this distinction is clear. In practice, a business is usually choosing between:
Ready-made SaaS agent: a pre-built product, often priced per seat, credit or resolution, with little or no conventional build cost.
Configured SaaS implementation: an existing platform's agent feature, set up and connected to a company's own data and systems.
Low-code agent workflow: built using a platform's visual tools, with moderate setup effort and limited custom development.
Bespoke agent built on model APIs: designed and coded from the ground up, with substantial design, integration, security and evaluation work. Multi-step workflows of this kind, connecting several business systems together, are covered in more general terms in our guide to AI agents for small businesses.
Enterprise orchestration platform: governs several agents and workflows together, typically with contracted capacity and compliance controls.
Internal development: built and maintained by existing staff, where cost sits mostly in time rather than vendor fees.
A ready-made customer-service product can have almost no build cost at all. A custom system spanning several business platforms can carry substantial design, integration, security and testing costs before it ever answers a live enquiry. The price bands below cannot be interpreted sensibly without knowing which of these six a quote actually describes.
Ready-Made SaaS (lowest cost) → Configured Platform → Bespoke Agent → Enterprise Orchestration (highest cost)
Purchasing route | Upfront cost | Recurring basis | Best suited to | Main risk |
|---|---|---|---|---|
Ready-made SaaS | Low | Seat, credit or resolution | Standard, common workflow | Limited customisation |
Configured platform | Medium | Subscription plus usage | Existing integrations | Implementation creep |
Bespoke agent | High | API, hosting and support | Differentiated workflow | Ongoing maintenance burden |
Enterprise orchestration | Highest | Contracted capacity | Multiple governed workflows | Complexity and vendor lock-in |
AI Workforce planning estimate, August 2026: these are indicative planning bands based on typical project scope, not independently verified market averages, and should be treated as a starting point for scoping a conversation with a vendor rather than a quote.
A narrowly scoped agent, covering one channel, one approved information source and limited integration work, typically sits in a low five-figure sterling budget: roughly £6,000 to £24,000 for implementation. A multi-step production workflow involving several connected systems, consequential actions, custom security controls or extensive testing, the kind of data-protection and permissioning work covered in our AI and GDPR Compliance guide, typically moves into six figures, commonly £40,000 to £120,000, with enterprise-grade orchestration and compliance work sometimes exceeding £155,000.
Running costs are separate from build cost and are just as easy to underbudget. For initial planning, AI Workforce uses an indicative monthly running-cost band of approximately £150 to £2,000 or more for smaller deployments, depending on volume, tools and oversight. AI Workforce also uses an indicative annual maintenance allowance of approximately 15% to 30% of the original implementation cost where ongoing integration, evaluation and workflow changes are expected. This is a planning assumption, not an independently verified market average.
These bands are derived from AI Workforce's internal scoping assumptions for narrowly defined UK business deployments and should be replaced by a workflow-specific quotation before approval. Because they are planning assumptions rather than sourced market data, the reliable way to price an actual deployment is to estimate its integrations, actions, volumes, oversight and service requirements individually, using a structured framework such as the one below.
AI Workforce developed the AI Workforce Agent Cost Model as a practical framework for scoping and interpreting an AI agent quotation, rather than relying on a single headline price.
Scope → Systems → Data → Actions → Volume → Oversight → Service Level → Change
Scope: how many workflows and exception types does the agent need to cover?
Systems: which platforms, such as a CRM, booking system or payment provider, must it connect to?
Data: how much preparation, permissioning and retrieval-content work is required before launch?
Actions: does the agent only answer questions, or does it recommend, act or transact?
Volume: how many conversations, tasks, minutes or tokens will it realistically handle each month?
Oversight: how much human review and exception handling does the workflow need?
Service level: what uptime, latency and support commitments does the business require?
Change: how often will policies, connected systems and knowledge sources need to be updated?
A quote that only addresses volume and model choice, without addressing the other six factors, is unlikely to reflect the real total cost of ownership.
First-year total cost of ownership = implementation + licences + model and tool usage + hosting + monitoring + human review + maintenance + contingency
Unit cost = first-year total cost of ownership ÷ verified successful outcomes
The components typically stack up in this order of weight, though actual proportions vary by deployment: implementation, model and tool usage, human review, maintenance, monitoring, licences, hosting, contingency.
A "successful outcome" needs to be defined in advance, whether that is a resolved enquiry, a qualified lead or a completed task, so that unit cost can actually be measured after launch rather than estimated in the abstract.
For agents built directly on model APIs, charges normally distinguish input tokens, cached input tokens and output tokens, and rates differ materially by model and processing tier. The table below shows OpenAI's current published API pricing, checked August 2026, converted to sterling at the rate stated above:
Model checked | Input / 1M tokens | Cached input / 1M tokens | Output / 1M tokens | Checked |
|---|---|---|---|---|
GPT-5.6 Luna | $0.20 (£0.16) | $0.02 (£0.02) | $1.20 (£0.95) | August 2026 |
GPT-5.6 Terra | $2.00 (£1.58) | $0.20 (£0.16) | $12.00 (£9.48) | August 2026 |
GPT-5.6 Sol | $5.00 (£3.95) | $0.50 (£0.40) | $30.00 (£23.70) | August 2026 |
GPT-5.6 Luna is positioned as a fast, affordable model for high-volume everyday work. GPT-5.6 Sol is its flagship model, positioned for more demanding agentic work. Cached input, for repeated content such as a system prompt or reference document, is priced around 90% lower than standard input across all three current tiers, though this applies only to eligible repeated input, not to output, tool calls or the deployment's total running cost. Because model names and prices change frequently, always check OpenAI's current pricing page directly before budgeting against a specific model.
Two techniques are now standard practice for reducing this spend. Batch processing, for work that does not need an instant response, is priced at a 50% discount on standard input and output rates through OpenAI's Batch API. Prompt caching reduces the cost of repeated input, as above, but neither technique touches the cost of tool calls, hosting, monitoring or human review, which sit outside the model bill entirely. Not every AI agent product is priced this way: some bill by seat, resolution, conversation, credit or minute instead of by token, so the first step in any comparison is establishing which unit a quote is actually charging for.
Vendors structure their offers around four broad approaches, and the differences matter for what a business ends up paying. A flat fee per seat per month is predictable but can be poor value if usage is uneven. Billing tied directly to model and tool usage scales precisely with demand but is harder to forecast. A fixed charge for each completed action sits between the two. Outcome-based pricing, the newest and fastest-growing structure, charges only when the system successfully completes a defined task.
This structure is now reshaping how established vendors charge for their agent products. From 14 April 2026, HubSpot moved its Breeze Customer Agent from a flat per-conversation charge to £0.40 per resolved conversation (converted from a quoted $0.50, billed via HubSpot Credits), and its Breeze Prospecting Agent to £0.79 per recommended lead (converted from a quoted $1.00). Zendesk uses a similar principle with its "automated resolution" billing unit, an approach directly relevant to the containment and cost-per-resolved-contact metrics covered in our AI Call Centre guide: an AI agent that only assists before a human resolves the conversation is not charged against a customer's resolution allowance, while a "verified resolution," confirmed by a follow-up check that the customer did not need further help, is the unit that counts toward billing. Hybrid structures combining a baseline subscription with usage-based billing are increasingly common, so the most useful comparison is cost per successful outcome against realistic expected volume, since a low headline price can still produce a large bill at scale.
Unadvertised line items are one of the most common reasons budgets run over. Beyond model usage and tool charges, businesses routinely underestimate data preparation, ongoing monitoring, human review and the engineering time needed when a deployment starts failing on edge cases it was never tested against. Total operating cost may not scale in direct proportion to conversation volume, because individual requests can differ significantly in length, tool use, retries, reasoning steps and how much human review they require.
Ongoing support is another frequently overlooked item. Someone needs to review flagged exceptions and keep reference content and instructions current as products or policies change, including model and tool usage, hosting where applicable, monitoring, evaluation, knowledge-base updates and periodic workflow changes, rather than the periodic model retraining that many people assume is required. Every time a connected system changes its own structure, that integration may need rebuilding. A budget that excludes monitoring, corrections and integration maintenance is likely to understate the true cost of running the agent.
There is no defensible universal percentage, and any figure presented as a fixed answer should be treated with caution. PagerDuty's 2025 survey of 1,000 IT and business executives at companies with at least roughly £395 million (about $500 million) in annual revenue, across the US, UK, Australia and Japan, found that respondents expected an average 171% return on agentic AI, with 62% expecting a return above 100%. That figure is a forecast reported by senior executives, not an independently verified, realised return.
BCG's AI Radar 2026, a global survey of more than 2,300 business leaders including 640 CEOs across 16 markets, found that about 90% of CEOs believed AI agents would produce measurable returns during 2026, and around 80% said they felt more optimistic about AI's potential return than a year earlier. Both figures show strong confidence among senior leaders. Neither guarantees that any individual deployment will pay back its cost, and some AI projects are cancelled after launch specifically because expected returns and success criteria were never clearly defined before the build began.
Illustrative worked example, not vendor or company data.
A business handling 5,000 customer enquiries a month finds that 3,000 are suitable for automation, and an AI agent successfully resolves 2,250 of those without escalation.
5,000 enquiries → 3,000 eligible → 2,250 resolved → £4,000 monthly cost → £1.78 per successful outcome
At an illustrative outcome price of £0.40 per resolved conversation, resolution charges come to £900 a month. A platform subscription minimum adds £250. Reviewing the roughly 750 enquiries a month that still need a person, at an average of six minutes each and a fully loaded staff cost of £18 an hour, adds around £1,350. Implementation cost of £18,000 is allocated evenly across the first 12 months as a management-accounting choice; a business may use a different accounting or evaluation period, and this allocation adds a further £1,500 a month here. Total monthly cost: approximately £4,000, against 2,250 successful resolutions, giving a cost per successful outcome of roughly £1.78.
For simplicity, this example assumes hosting and basic monitoring are included in the £250 platform charge, and it excludes separate maintenance and contingency. A real first-year total cost of ownership calculation should add those items where applicable. If the business's verified baseline cost for resolving the same eligible enquiry types through the existing human-led process was around £3.20, this illustrative deployment would represent a saving of close to 44% per resolved enquiry, once implementation, subscription and ongoing human review are all included, not just the vendor's per-resolution price. The exercise matters more than the specific numbers: any real evaluation should run through the same steps using a company's own volumes, escalation rate, staff cost and vendor pricing.
AI Workforce developed the AI Workforce Agent Boundary Matrix to help businesses separate cost that is easy to forecast from cost that depends on how the deployment actually performs.
Predictable: subscriptions, committed usage allowances and fixed platform fees.
Volume-sensitive: tokens, minutes, searches and individual tool calls, which scale directly with usage.
Operationally variable: human review, exception handling and customer support, which depend on how well the agent performs in practice.
Change-driven: integration maintenance, policy updates, evaluation work and periodic redesign as the business and its systems evolve.
Voice deployments also incur telephony, transcription and generated-audio costs on top of model usage, so businesses evaluating a phone-based agent should model cost per completed call separately; our AI Voice Agents guide covers this in more depth. A budget that only accounts for the predictable and volume-sensitive categories will consistently understate what a deployment actually costs to run.
Work Out What Your AI Agent Would Actually Cost. Use the AI Readiness Assessment to identify the workflow, integrations, oversight and usage assumptions that should go into a realistic budget.
Cost control starts with matching the model to the task rather than defaulting to the most capable option available. A system answering routine questions rarely needs a flagship model; a smaller, cheaper model handling large-volume, low-complexity requests can cut spend meaningfully while barely affecting quality. Routing that sends straightforward requests to a cheaper model, and only escalates genuinely difficult ones to a more capable model, is now standard practice among cost-conscious teams.
Caching, batching and lean prompt design are the next lever, and none of them touch implementation cost at all. The final lever is scope discipline: businesses that prove one well-defined workflow first, using the kind of structured self-check covered in our AI Readiness Assessment, and measure the outcome against a specific, agreed definition of success before expanding, are better placed to control spend, identify correction costs and decide whether expansion is justified.
Cost does not scale in a straight line between something simple and something complex; it scales with the number of steps, integrations and decision points involved. A narrow, single-job agent, such as one answering from a fixed knowledge base, is comparatively cheap to build and run because each request typically involves only a small number of model and tool calls and minimal orchestration. A system that retrieves information, calls multiple external systems, reasons across several steps and hands off to a person when uncertain adds cost at every one of those layers.
This is why, depending on task complexity, the same underlying language models can produce very different bills. Large-volume, simple jobs benefit most from aggressive optimisation and cheaper models, since even small per-request savings compound quickly at scale. Complex, multi-step deployments are harder to optimise on unit cost alone; the more effective lever is usually removing unnecessary steps from the workflow itself, since every additional call adds both latency and cost.
What determines the cost of an AI agent?
Cost is driven far more by the workflow around the model than by the model itself: the number of integrations, how much data preparation and permissioning is required, whether the agent only answers or also acts, expected volume, how much human oversight it needs, and how often connected systems and policies change.
What is included in AI-agent running costs?
Running costs typically cover model and tool usage, hosting where applicable, monitoring, human review of flagged exceptions, and periodic updates to reference content, instructions and integrations as the business changes.
How do token and outcome-based prices differ?
Token-based pricing charges for the model's input and output regardless of whether the task succeeds, and is transparent for high-volume, well-understood workloads. Outcome-based pricing charges only when a task is completed successfully, shifting more risk to the vendor, which suits less predictable volume.
What is the highest hidden cost?
Integration and ongoing maintenance are often among the most underestimated costs. Connecting an agent to live business systems, and keeping those connections working as those systems change, can cost more over time than the underlying model usage.
How long does it take to see a return?
There is no reliable universal payback period. It depends on implementation cost, adoption, transaction volume, the value of each successful outcome and the cost of human review. Establish the current cost of the workflow before launch, then calculate payback from verified savings or added contribution, not from expected activity or industry-wide averages.
Is a ready-made agent cheaper than a bespoke one?
Often, for a narrow, common job. An off-the-shelf agent can get a business running for a few hundred pounds a month, but a genuinely custom, production-grade system that spans several internal platforms will usually cost more to build and maintain properly, and that additional cost can still be worthwhile if the workflow is differentiated enough to justify it.
AI agent cost is not one number: it combines implementation spend, ongoing running costs and often-overlooked maintenance costs, and depends heavily on which of the six purchasing routes a quote actually describes
AI Workforce's planning bands are roughly £6,000 to £24,000 for a narrowly scoped agent, and roughly £40,000 to £120,000 or more for a complex, multi-system deployment; treat these as indicative starting points, not fixed market prices
Token and API costs scale with volume and complexity, and cached input and batch processing can meaningfully reduce eligible spend, but neither touches tool calls, hosting or human review
Vendors charge through per-seat, per-token, per-task or outcome-based pricing, each shifting risk differently between vendor and customer, and established platforms including HubSpot and Zendesk have moved toward outcome-based billing
PagerDuty's 2025 survey found executives expected an average 171% return from agentic AI, and BCG's AI Radar 2026 found about 90% of CEOs believed AI agents would deliver measurable returns in 2026; both are expectations, not verified realised returns
There is no reliable universal payback period; calculate it from a company's own verified savings, not from an industry-wide average
The most useful way to compare vendors is total cost of ownership divided by verified successful outcomes, not build price or token cost in isolation
The safest way to control spend is to prove one well-defined workflow first, define success in advance, and expand only once it is confirmed
PagerDuty, "The Next Generation of AI: More than Half of Companies (51%) Already Deployed AI Agents": 2025 Agentic AI Survey of 1,000 IT and business executives; the 171% figure is a reported expectation, not a verified realised return.
Boston Consulting Group, "CEOs Are Taking Charge of AI," AI Radar 2026 Weekly Brief, January 2026: survey of 2,300+ business leaders including 640 CEOs; the 90% figure reflects CEO belief, not measured outcomes.
HubSpot, "HubSpot's Customer Agent and Prospecting Agent: Now You Pay When the Task Is Complete," April 2026: Breeze Customer Agent and Prospecting Agent outcome pricing, effective 14 April 2026.
Zendesk, "About Automated Resolution Tiers," Zendesk Help Centre: describes the resolution-tier billing structure referenced above.
OpenAI, "API Pricing," checked August 2026: current GPT-5.6 Luna, Terra and Sol pricing for input, cached input and output tokens, and Batch API discount.
AI Workforce internal planning estimates for build cost, running cost and maintenance bands, August 2026: indicative figures based on typical project scope, not independently verified third-party market data.
All dollar-denominated figures are converted to sterling at an illustrative rate of $1.27 to £1 (August 2026) and should be treated as indicative rather than exact; check a vendor's current pricing page directly before budgeting.
AI Workforce can scope the implementation, integrations, expected usage, human oversight and first-year total cost for your proposed agent.
Luca Controlo is AI Adoption and Marketing Automation Lead at AI Workforce, a British AI company building AI agents for UK businesses. He works with UK businesses to scope AI agent projects against realistic budgets and measurable outcomes.
Reviewed by Seth Ayush, Co-Founder of AI Workforce · August 2026