AI cost and economics atlas

Price the workload.
Measure the system.

A practical decision system for API pricing, self-hosted inference, GPU capacity, utilisation and the operating costs hidden between a token rate and a dependable service.

Precision balance connecting a metered stream of tokens with a cooled GPU infrastructure system
DEMAND / TOKENS / CAPACITY / OPERATIONSCST-01
20PRICED API ENDPOINTS
09GPU COST PRESETS
API, dedicated endpoint and self-hosted Official provider and infrastructure sources Last researched 27 Jul 2026

Total cost of inference

The invoice is only
one layer.

Token rates explain an external API bill. They do not explain quality, retries, context growth, latency, idle capacity, staffing, observability, resilience or the cost of changing direction. A useful comparison keeps workload, service target and operating model visible at the same time.

Exploded AI economics stack connecting demand, model, serving, compute, cooling and operations layers
FROM DEMAND TO OPERATING REALITYSTACK-06

ECONOMICS MAP / 06 CONNECTED LAYERS

Count every layer.
Keep each assumption.

  1. 06
    Business outcomeCompletion quality, time saved, revenue protected and risk controlled
  2. 05
    Demand shapeRequests, peaks, input length, output length, retries and growth
  3. 04
    Model and routingCapability tier, reasoning, cache policy, fallbacks and tool loops
  4. 03
    Serving systemBatching, quantisation, context, throughput, latency and replicas
  5. 02
    InfrastructureGPU hours, memory, network, storage, power and facility
  6. 01
    Operations and riskEngineering, observability, support, availability, security and change

Six cost levers

Optimise the system.
Not one rate.

The largest savings usually come from choosing the right model for each task, controlling output, reusing stable context and measuring the complete production loop.

01Route

Send routine work to an efficient model and reserve premium reasoning for tasks that earn it.

02Cache

Stabilise repeated prompt prefixes, then distinguish cache reads, cache writes and cache storage.

03Bound

Limit unnecessary context, reasoning, tool loops and output while preserving task quality.

04Batch

Move delay-tolerant work to provider batch tiers or fuller local batches when the service permits.

05Measure

Track successful task cost, latency, retries and discarded output, not tokens in isolation.

06Operate

Price redundancy, monitoring, security, upgrades and engineering ownership before self-hosting.

Interactive scenario model

Find the economic
break-even point.

Enter a real workload, select a current API endpoint and test it against a priced GPU deployment. The result is a planning model, not a quote or benchmark. Throughput must come from a workload-matched measurement at the required latency target.

01 / WORKLOAD

Demand and API

Monthly requests
0
Input tokens
0
Output tokens
0
02 / INFRASTRUCTURE

Capacity and operations

03 / RESULT

Monthly comparison

API ESTIMATE $0 $0 / 1M output
SELF-HOSTED ESTIMATE $0 $0 / 1M output
API
SELF

Ordinary input
$0
Cached input
$0
API output
$0
GPU capacity
$0
Host overhead
$0
Capacity coverage
0x
API break-even demand 0 / day At the entered token ratio and current standard API rates

The self-hosted path assumes a separately selected deployable model can meet the same task-quality target. It is not the selected proprietary API model. This calculator excludes API tool charges, cache-write premiums, cache storage, taxes, failed requests, quality differences, migrations and contract discounts. Cloud node prices already include provider facility power. Do not add electricity or PUE to those hourly presets.

Current API price directory

Compare the billable
units first.

Twenty high-interest production and preview endpoints are shown at their standard synchronous text rates. Input, cached input and output are separate because the cheapest line item is rarely the whole bill. Capability and price are not interchangeable.

20 endpoints shown Prices in USD per 1M text tokens Research snapshot: 27 Jul 2026

API and self-hosted

The crossing point
moves.

Self-hosting can become attractive when demand is steady, measured throughput is high and the team can keep expensive capacity productive. It becomes weak when traffic is uncertain, replicas sit idle or the service needs capabilities a smaller model cannot deliver.

01 / VOLUME
Higher sustained demand spreads a fixed infrastructure floor across more successful work.
02 / UTILISATION
Paid capacity that waits for traffic still appears in the monthly total.
03 / QUALITY
A cheaper model is not cheaper when lower task success creates retries or human rework.
04 / RELIABILITY
A production comparison needs independent failure domains, observability and an upgrade path.
API token metering and self-hosted GPU infrastructure converging through a transparent break-even chamber
PAY PER USEBREAK-EVENOWN CAPACITY

Deployment economics

Six operating models.
Six different floors.

Choose the service model before comparing prices. A token API, managed endpoint, single node and rack-scale fleet transfer different costs and responsibilities to the buyer.

01 / EXTERNAL API

No idle GPU floor

Best when demand is uncertain, bursty or changing quickly. Pay for billable use and retain access to provider-managed model and infrastructure updates.

WATCH / RATE LIMITS, DATA CONTROLS, TOOL FEES, PROVIDER DEPENDENCY
02 / MANAGED ENDPOINT

Dedicated without the platform team

Useful when a private model endpoint matters but the team does not want to own the complete serving stack. Minimum replicas and initialization time create a floor.

WATCH / SCALE-TO-ZERO, COLD STARTS, MINIMUM REPLICAS, REGION
03 / SINGLE GPU

Low entry cost, one failure domain

A strong development and internal-workload pattern for smaller or quantised models. It is not a high-availability production design by itself.

WATCH / MODEL FIT, KV CACHE, MAINTENANCE, OUTAGE WINDOW
04 / DUAL INDEPENDENT NODES

A practical production baseline

Two independently hosted nodes can survive a host failure and support safer upgrades. The cost floor is at least twice the single-node rate before operations.

WATCH / LOAD BALANCING, STATE, FAILOVER, CAPACITY HEADROOM
05 / EIGHT-GPU NODE

Large model, high idle floor

Model parallelism, large context and sustained throughput can justify the node. Interconnect, serving engine and workload shape decide whether the system is productive.

WATCH / FABRIC, BATCHING, MEMORY, CONTINUOUS UTILISATION
06 / RACK-SCALE FLEET

Infrastructure becomes a product

Suitable only for very large or continuously utilised workloads. Networking, power, cooling, spares, orchestration and specialist operations become first-class costs.

WATCH / FACILITY, NETWORK, CAPACITY PLANNING, OWNERSHIP

Research and calculation method

Make the assumptions
auditable.

Every price is linked to the current official provider page. The calculator keeps provider rates, workload volume, measured throughput and operational overhead separate so a change in one assumption can be seen and tested.

01

Use standard rates

The directory starts with synchronous public list pricing. Contract, batch, priority, flex, regional and tool charges remain separate.

02

Separate token classes

Ordinary input, cache reads, cache writes, output and reasoning can carry different rates and rules.

03

Measure at the SLO

Self-hosted throughput must be measured at the required latency, context, model, precision, concurrency and quality.

04

Price productive output

Retries, discarded answers and test traffic consume capacity but do not create successful production output.

05

Keep service models distinct

Full cloud VMs, GPU-only listings and managed endpoints contain different CPU, memory, network and support value.

06

Recheck before purchase

Pricing, endpoint status, capacity and model terms change. Confirm every current source before committing spend.

Before choosing a cost path

Six questions before
the forecast.

01

What is a successful unit?

Price a completed task, approved document, resolved case or productive session, not a token in isolation.

02

How does demand peak?

Average volume can hide the concurrency and spare capacity needed during a short operational peak.

03

Which model earns its tier?

Route by measured task quality and reserve expensive reasoning for decisions that need it.

04

What can be reused?

Stable instructions, reference material and repeated context may benefit from caching or retrieval.

05

Who owns the service?

Include deployment, security, monitoring, incident response, upgrades and model evaluation in the operating model.

06

What changes the answer?

Run sensitivity cases for growth, utilisation, output length, model price, contract discount and redundancy.

Engineering the economic system

From workload model
to production reality.

ADOR.IS designs AI-enabled products, model-routing systems, inference infrastructure and operational software around measurable service and cost constraints.

[email protected]