Send routine work to an efficient model and reserve premium reasoning for tasks that earn it.
AI cost and economics atlas
Price the workload.
Measure the system.
A practical decision system for API pricing, self-hosted inference, GPU capacity, utilisation and the operating costs hidden between a token rate and a dependable service.
Total cost of inference
The invoice is only
one layer.
Token rates explain an external API bill. They do not explain quality, retries, context growth, latency, idle capacity, staffing, observability, resilience or the cost of changing direction. A useful comparison keeps workload, service target and operating model visible at the same time.
ECONOMICS MAP / 06 CONNECTED LAYERS
Count every layer.
Keep each assumption.
- 06Business outcomeCompletion quality, time saved, revenue protected and risk controlled
- 05Demand shapeRequests, peaks, input length, output length, retries and growth
- 04Model and routingCapability tier, reasoning, cache policy, fallbacks and tool loops
- 03Serving systemBatching, quantisation, context, throughput, latency and replicas
- 02InfrastructureGPU hours, memory, network, storage, power and facility
- 01Operations and riskEngineering, observability, support, availability, security and change
Six cost levers
Optimise the system.
Not one rate.
The largest savings usually come from choosing the right model for each task, controlling output, reusing stable context and measuring the complete production loop.
Stabilise repeated prompt prefixes, then distinguish cache reads, cache writes and cache storage.
Limit unnecessary context, reasoning, tool loops and output while preserving task quality.
Move delay-tolerant work to provider batch tiers or fuller local batches when the service permits.
Track successful task cost, latency, retries and discarded output, not tokens in isolation.
Price redundancy, monitoring, security, upgrades and engineering ownership before self-hosting.
Interactive scenario model
Find the economic
break-even point.
Enter a real workload, select a current API endpoint and test it against a priced GPU deployment. The result is a planning model, not a quote or benchmark. Throughput must come from a workload-matched measurement at the required latency target.
The self-hosted path assumes a separately selected deployable model can meet the same task-quality target. It is not the selected proprietary API model. This calculator excludes API tool charges, cache-write premiums, cache storage, taxes, failed requests, quality differences, migrations and contract discounts. Cloud node prices already include provider facility power. Do not add electricity or PUE to those hourly presets.
Current API price directory
Compare the billable
units first.
Twenty high-interest production and preview endpoints are shown at their standard synchronous text rates. Input, cached input and output are separate because the cheapest line item is rarely the whole bill. Capability and price are not interchangeable.
Try another provider, price band, context range or search term.
API and self-hosted
The crossing point
moves.
Self-hosting can become attractive when demand is steady, measured throughput is high and the team can keep expensive capacity productive. It becomes weak when traffic is uncertain, replicas sit idle or the service needs capabilities a smaller model cannot deliver.
- 01 / VOLUME
- Higher sustained demand spreads a fixed infrastructure floor across more successful work.
- 02 / UTILISATION
- Paid capacity that waits for traffic still appears in the monthly total.
- 03 / QUALITY
- A cheaper model is not cheaper when lower task success creates retries or human rework.
- 04 / RELIABILITY
- A production comparison needs independent failure domains, observability and an upgrade path.
Deployment economics
Six operating models.
Six different floors.
Choose the service model before comparing prices. A token API, managed endpoint, single node and rack-scale fleet transfer different costs and responsibilities to the buyer.
No idle GPU floor
Best when demand is uncertain, bursty or changing quickly. Pay for billable use and retain access to provider-managed model and infrastructure updates.
WATCH / RATE LIMITS, DATA CONTROLS, TOOL FEES, PROVIDER DEPENDENCYDedicated without the platform team
Useful when a private model endpoint matters but the team does not want to own the complete serving stack. Minimum replicas and initialization time create a floor.
WATCH / SCALE-TO-ZERO, COLD STARTS, MINIMUM REPLICAS, REGIONLow entry cost, one failure domain
A strong development and internal-workload pattern for smaller or quantised models. It is not a high-availability production design by itself.
WATCH / MODEL FIT, KV CACHE, MAINTENANCE, OUTAGE WINDOWA practical production baseline
Two independently hosted nodes can survive a host failure and support safer upgrades. The cost floor is at least twice the single-node rate before operations.
WATCH / LOAD BALANCING, STATE, FAILOVER, CAPACITY HEADROOMLarge model, high idle floor
Model parallelism, large context and sustained throughput can justify the node. Interconnect, serving engine and workload shape decide whether the system is productive.
WATCH / FABRIC, BATCHING, MEMORY, CONTINUOUS UTILISATIONInfrastructure becomes a product
Suitable only for very large or continuously utilised workloads. Networking, power, cooling, spares, orchestration and specialist operations become first-class costs.
WATCH / FACILITY, NETWORK, CAPACITY PLANNING, OWNERSHIPResearch and calculation method
Make the assumptions
auditable.
Every price is linked to the current official provider page. The calculator keeps provider rates, workload volume, measured throughput and operational overhead separate so a change in one assumption can be seen and tested.
Use standard rates
The directory starts with synchronous public list pricing. Contract, batch, priority, flex, regional and tool charges remain separate.
Separate token classes
Ordinary input, cache reads, cache writes, output and reasoning can carry different rates and rules.
Measure at the SLO
Self-hosted throughput must be measured at the required latency, context, model, precision, concurrency and quality.
Price productive output
Retries, discarded answers and test traffic consume capacity but do not create successful production output.
Keep service models distinct
Full cloud VMs, GPU-only listings and managed endpoints contain different CPU, memory, network and support value.
Recheck before purchase
Pricing, endpoint status, capacity and model terms change. Confirm every current source before committing spend.
PRIMARY SOURCE LEDGER / VERIFIED 27 JUL 2026
Before choosing a cost path
Six questions before
the forecast.
What is a successful unit?
Price a completed task, approved document, resolved case or productive session, not a token in isolation.
How does demand peak?
Average volume can hide the concurrency and spare capacity needed during a short operational peak.
Which model earns its tier?
Route by measured task quality and reserve expensive reasoning for decisions that need it.
What can be reused?
Stable instructions, reference material and repeated context may benefit from caching or retrieval.
Who owns the service?
Include deployment, security, monitoring, incident response, upgrades and model evaluation in the operating model.
What changes the answer?
Run sensitivity cases for growth, utilisation, output length, model price, contract discount and redundancy.
Engineering the economic system
From workload model
to production reality.
ADOR.IS designs AI-enabled products, model-routing systems, inference infrastructure and operational software around measurable service and cost constraints.
[email protected] ↗