The factors that drive the cost of building and running an AI chatbot or AI agent: scope, data, integrations, model usage, evaluation, security and ongoing improvement.
In this article
- 01Why AI project quotes vary so much
- 02Chatbot or agent: scope is the biggest factor
- 03Knowledge sources and data preparation
- 04Integrations with your systems
- 05Model choice and running costs
- 06Evaluation, safety and security
- 07Channels and user experience
- 08Ongoing improvement and support
- 09Build, buy or customize?
- 10How to get a realistic estimate
Why AI project quotes vary so much
Ask several vendors to price an AI chatbot or AI agent and the answers will differ widely. One may quote a simple FAQ bot connected to a few documents, while another assumes a multilingual assistant integrated with your CRM, order system and support desk, with detailed evaluation and security reviews. Both are called AI chatbots, but they are very different projects.
This guide explains the factors that drive cost, both for the initial build and for running the system afterward, so you can define scope clearly and compare quotes fairly. It does not list prices, because those depend on your specific requirements, but it shows which decisions move the budget most.
Chatbot or agent: scope is the biggest factor
A chatbot answers questions in conversation. An AI agent takes actions: it plans steps, calls tools and updates systems, such as processing a refund or booking an appointment. Agents require more design, integration, testing and safety controls, because mistakes affect real records rather than just producing a wrong answer. Our comparison of an AI agent vs a chatbot explains the difference in detail.
Within either category, scope depends on how many topics, workflows and user types the system must handle. A focused assistant for one task is far cheaper to build and test than a general assistant expected to handle anything.
- Number of topics, intents or workflows covered.
- Answer-only behavior versus actions in other systems.
- Number of user types with different permissions.
- Languages and channels supported.
Knowledge sources and data preparation
Most business assistants answer from company knowledge using retrieval-augmented generation, which searches your documents and passes relevant passages to the language model. The effort depends on how much content exists, how clean it is and how often it changes. Outdated, duplicated or scanned documents need cleanup before they produce reliable answers.
Permissions add work too. If different users may only see certain documents, retrieval must respect those rules, which requires integrating with document systems and testing access carefully. Multilingual content and frequently changing sources, such as price lists, add further ingestion and testing work.
Integrations with your systems
Every system the assistant connects to adds design, development and testing effort. Reading order status from an ecommerce platform, creating tickets in a help desk, updating CRM records or checking appointment availability each require secure access, error handling and clear rules about what the assistant may do.
Older systems without modern APIs increase effort further, sometimes requiring middleware or workarounds. Integration quality often determines whether an assistant genuinely reduces workload or simply hands most conversations back to staff. Read-only integrations are simpler than those that change records. Standard protocols such as the Model Context Protocol (MCP) make tool connections more reusable, but each one still needs its own permissions, error handling and testing.
- CRM, help desk and ticketing systems.
- Order, booking and inventory systems.
- Identity systems for user authentication.
- Messaging channels such as WhatsApp and web chat.
Model choice and running costs
Language model usage is usually billed per token, so running costs grow with conversation volume, prompt length and the size of retrieved context. Larger models cost more per request but handle complex reasoning better. Reasoning models also bill for the thinking tokens they use before answering. Many systems route simple questions to smaller, cheaper models, reserve larger models and longer reasoning for difficult requests, and use prompt caching to cut the cost of repeated instructions and context.
Other running costs include vector databases, hosting, monitoring, channel fees such as WhatsApp conversation charges and support time. Estimate expected volumes early, because a design that is affordable at a few hundred conversations a day can become expensive at many thousands.
Evaluation, safety and security
Reliable AI systems need evaluation: a set of real questions and expected answers used to measure quality before launch and after every change. Building this set, reviewing results and tuning prompts and retrieval takes time, but it is what separates a convincing demo from a dependable product.
Safety and security add further effort, especially for agents. Guardrails against prompt injection, limits on actions, human approval for sensitive steps, logging for audits and protection of personal data are essential in production and should be included in any serious estimate.
- Evaluation sets and quality scoring.
- Guardrails and content policies.
- Approval steps for high-impact actions.
- Tracing of model calls, retrieval and tool actions in production.
- Audit logs and data protection controls.
Channels and user experience
Deploying on a website widget, inside a mobile app, on WhatsApp, in Microsoft Teams or over the phone each adds work. Voice assistants are particularly demanding because speech recognition, latency and natural conversation flow must work together. Handoff to human agents, conversation history and analytics also need design and integration.
Design work matters as much as the model. Clear greetings, suggested questions, sensible fallbacks when the assistant is unsure and a visible route to a person shape how customers judge the experience. Testing conversation flows with real users before launch reveals confusing moments that internal testing often misses.
Ongoing improvement and support
AI assistants are never finished. Content changes, users ask new questions and model providers release new versions. Budget for regular reviews of conversations, content updates, prompt improvements and re-evaluation after model upgrades. Teams that skip this maintenance usually see quality decline within months.
A practical approach is to start with a focused pilot, measure resolution rates and user satisfaction, then expand scope based on evidence. This keeps early investment small and directs further spending toward the improvements that matter most.
Build, buy or customize?
Not every assistant needs to be built from scratch. Many help desk, CRM and messaging platforms now include AI features that can answer common questions from your knowledge base with modest configuration. They are quick to launch and suit standard support needs, though customization, data control and integration depth may be limited.
Custom development makes sense when the assistant must follow your specific workflows, connect deeply with internal systems, meet strict data residency or security requirements, or deliver a distinctive customer experience. A hybrid approach is common: start with a platform feature to learn what customers ask, then build custom capabilities where the platform falls short.
- Built-in platform AI features for standard support questions.
- Configurable chatbot platforms for moderate customization.
- Custom assistants and agents for complex workflows and integrations.
How to get a realistic estimate
Prepare a short brief listing the questions or tasks the assistant must handle, the systems it must connect to, expected conversation volumes, languages and channels. Share sample documents and real customer questions. With this information, a team experienced in AI chatbot development or AI agent development can propose a phased plan and estimate. For broader software budgets, our app cost calculator offers a useful starting point.



