The useful answer to “How much does custom AI development cost?” is not a single market average. It is a map of the uncertainty your company must retire before a production decision becomes responsible. A document assistant connected to a curated knowledge base has a different cost structure from an agent that can modify customer records, and both differ from a model trained on proprietary data.
For US buyers, the budget usually reflects five connected workstreams: product discovery, data readiness, model or system engineering, integration, and production operations. Omitting one does not make it free. It moves the expense into rework, manual review, security remediation, or an operations team that inherits a system it cannot explain.
The four planning bands
A bounded feasibility engagement should answer whether the workflow is technically plausible and what evidence would justify further investment. A pilot should connect a thin vertical slice to representative data and users. A production build adds identity, permissions, integrations, observability, failure handling, and support ownership. A continuing optimization program maintains evaluations as models, prompts, data, and business rules change.
These bands are more useful than a universal price because they connect spending to a decision. A buyer can fund the smallest phase that produces the next credible piece of evidence.
- Feasibility: workflow definition, baseline, data sample, evaluation design, and a go/revise/stop recommendation.
- Pilot: a usable end-to-end slice, limited integration, representative testing, and operating feedback.
- Production: security, reliability, monitoring, accessibility, documentation, rollout, and handover.
- Optimization: regression evaluation, model or provider changes, cost control, and workflow expansion.
What changes the cost most
The largest driver is rarely the number of screens. It is the consequence of being wrong. A low-risk drafting tool can tolerate review and occasional failure. A system that approves payments, changes inventory, or communicates regulated advice needs tighter authorization, evidence, auditability, and human control.
Data condition also changes the plan. Clean, permissioned, well-labeled information can support an early baseline. Fragmented documents, ambiguous ownership, sensitive fields, or missing ground truth require data engineering and governance before model quality can be judged honestly.
- Number and volatility of systems being integrated
- Required quality, latency, availability, and auditability
- Human-review and exception-management workflow
- Volume, sensitivity, and readiness of enterprise data
- Whether the model is hosted, open-weight, adapted, or trained
- Internal team capacity for product and operational ownership
A timeline should be evidence-gated
A calendar-only roadmap assumes the unknowns will cooperate. An evidence-gated roadmap defines what must be true before the team expands scope. For example: retrieval quality must cross an agreed threshold before interface polish; tool permissions must pass abuse tests before autonomous execution; the owning team must accept runbooks before broad rollout.
This approach protects speed. Weak paths stop early, while strong paths receive investment with less debate. The result is not a promise that AI research is predictable. It is a delivery system that makes uncertainty visible.
How to compare proposals
Ask each provider to separate assumptions from deliverables. A credible proposal identifies the target workflow, evaluation population, integration boundary, acceptance decision, security responsibilities, and the artifacts your team will own. Proposals that lead with a model name but do not define failure costs are incomplete.
Build the estimate from work packages
A defensible estimate separates discovery, data preparation, system engineering, application integration, evaluation, security, rollout, and ongoing operations. Each package should name its assumptions, decision owner, exit evidence, and excluded work. This makes proposals comparable even when vendors organize their teams differently. It also prevents a low initial estimate from hiding necessary production work in a later change request.
Use ranges while material uncertainty remains. Narrow the range after the team has inspected representative data, tested the most consequential integration, and established a measurable baseline. Precision before those activities is presentation, not planning. A responsible partner explains which experiment or discovery artifact will reduce each uncertainty and how the result changes the next funding decision.
Plan the lifetime operating cost
Production cost includes model inference, retrieval infrastructure, observability, human review, incident handling, evaluation maintenance, provider changes, and engineering ownership. These costs behave differently as usage grows. Token consumption may scale with requests, while evaluation and compliance work scale with changes, new workflows, and the consequence of failure.
Ask for a simple operating model at expected, high, and stress volumes. It should expose request mix, context size, latency targets, review rate, storage, and support assumptions. The purpose is not to predict every invoice. It is to identify which product decisions control cost and which costs remain after a model price falls.
Executive decision record
The decision is whether the proposed first phase retires enough uncertainty to justify a production investment, rather than whether a vendor can demonstrate an attractive prototype. Write that decision before selecting a model, vendor, framework, or implementation pattern. A written boundary keeps technical exploration connected to the operating outcome and makes it possible to explain why the organization advanced, revised, or stopped the work.
Approval should depend on a representative baseline, inspected data and integration boundaries, an evaluation design, explicit assumptions, and a phased estimate tied to acceptance decisions. The evidence does not need to remove every uncertainty, but it should address the uncertainty capable of changing value, architecture, risk, or ownership. Record the baseline, assumptions, unresolved questions, and the person accepting the next stage.
Failure boundary and operating ownership
The central failure to guard against is funding a calendar and feature list while data, authorization, evaluation, and operating responsibilities remain unresolved. Treat that condition as a testable scenario. Define how the system detects it, what users experience, which action is prevented or reversed, and what evidence reaches the person responsible for recovery.
Long-term accountability sits with a joint product and engineering sponsor who can accept scope tradeoffs and fund the controls required by the consequence of error. Supporting specialists can provide platforms, research, review, or delivery capacity, but they cannot substitute for an owner who controls policy and operating change. Name that owner before production and include the ownership path in release evidence and incident procedure.
A practical 90-day application plan
During the first 30 days, convert a representative baseline, inspected data and integration boundaries, an evaluation design, explicit assumptions, and a phased estimate tied to acceptance decisions into a bounded evidence plan. Assign each artifact to a named contributor, identify the representative inputs required, and agree on the comparison baseline before implementation expands. The objective of this period is to expose the assumption most likely to invalidate the work while the cost of changing direction is still low.
During days 31 through 60, build or instrument the smallest complete workflow that can support the decision about whether the proposed first phase retires enough uncertainty to justify a production investment, rather than whether a vendor can demonstrate an attractive prototype. Include the real data and authorization path where feasible, record exceptions, and review difficult cases with the people who own the underlying process. Resist adding breadth until the team can explain the measured behavior of this narrow slice.
During days 61 through 90, test the boundary represented by funding a calendar and feature list while data, authorization, evaluation, and operating responsibilities remain unresolved. Exercise degraded dependencies, ambiguous inputs, recovery, and handoff rather than demonstrating only successful cases. End the period with a written advance, revise, or stop decision that cites evidence, residual exposure, expected operating cost, and the next authority boundary.
The review should be accepted by a joint product and engineering sponsor who can accept scope tradeoffs and fund the controls required by the consequence of error. That group should confirm not only that the system can work, but that ownership, support capacity, monitoring, and change control are credible. If those conditions are absent, the responsible outcome is another bounded learning stage rather than an unsupported production commitment.
Practical checklist
- Define the business decision or workflow before selecting a model.
- Provide representative inputs and difficult failure examples.
- Agree on who can approve, override, and stop the system.
- Require evaluation, observability, documentation, and handover deliverables.
- Budget the production path, not only the demonstration.
- Use a bounded first phase with an explicit investment decision.
Engagement scenario
A US professional-services firm wants an assistant that prepares first drafts from a large internal document library. The first phase compares retrieval and long-context baselines on a permissioned sample, defines citation accuracy and omission tests, and maps document access rules. Only after the evidence supports the workflow does the program add identity integration, review queues, telemetry, and a controlled department rollout.
Continue reading
- [AI agent architecture for enterprise workflows](/enterprise-ai-agent-architecture)
- [An evaluation framework for production LLM systems](/llm-evaluation-framework-production)
- [How to prioritize an enterprise AI portfolio](/ai-use-case-prioritization-framework)
Sources and further reading
- [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework)
- [Google guidance on people-first content and evidence](https://developers.google.com/search/docs/fundamentals/creating-helpful-content)