Cost of querying a model
At GPT-3.5-level MMLU performance
Cost fell from about US$20 to US$0.07 per million tokens - more than a 280× decline. Source: Stanford HAI AI Index 2025.
Solveworx
The models are becoming cheaper. Poor AI engineering is what makes enterprise AI expensive.
Model prices are falling fast, yet many enterprise AI bills are rising.
At GPT-3.5-level MMLU performance
Cost fell from about US$20 to US$0.07 per million tokens - more than a 280× decline. Source: Stanford HAI AI Index 2025.
Lower unit prices can coincide with higher total spend as workloads, reasoning depth, context and agent systems grow.
Leading providers cluster on Arena-style ratings
When model access is less differentiating, optimisation, reliability and economics matter more. Source: Stanford HAI AI Index 2026.
Inference cost is influenced by decisions across the stack, not just the model or token price.
Common sources of avoidable cost in enterprise AI.
Using larger, more expensive models than needed.
Unnecessarily long context increases token usage and latency.
Poor retrieval quality produces more tokens and lower accuracy.
Iteration without guardrails can consume significant GPU time.
Low utilisation without pooling and dynamic scaling.
Inefficient settings, KV-cache, batching and routing increase latency.
Maximise business value per dollar spent.
Workloads, costs, quality, usage
Tokens, GPU utilisation, latency, cost/request
Select the best model and inference path
Tune batching, KV-cache and routing
Cost caps, policies and production readiness
Track improvement and business impact
Multiple layers to tune - from compute to business outcomes.
User experiences and business outcomes
Task decomposition, tool use and orchestration
High-quality, relevant context
Model routing, speculative decoding, KV-cache and batching
Model selection, quantisation and adapters
Efficient utilisation, pooling and autoscaling
An AI optimisation service: expertise and platform working together.
AI science & optimisation
Orchestration platform
A governed, optimised path to measurable outcomes.
Unrestricted experimentation can drive unpredictable costs.
Policies, model routing and context discipline.
Right model, right settings, efficient serving.
Fixed monthly pricing and cost caps.
Higher productivity, better service, clear ROI.
Stanford HAI AI Index 2025Trends in AI model pricing and inference costs.
Stanford HAI AI Index 2026Model performance convergence across leading providers.
FinOps FoundationOptimizing GenAI Usage: industry guidance on inference spend and GPU utilisation.
Xia et al. (2024)Unlocking Efficiency in Large Language Model Inference.
Jiang et al. (2026)Towards Efficient Large Language Model Serving.
Li et al. (2024)A Survey on Large Language Model Acceleration based on KV Cache.