Use the precision and sparsity mode the workload can sustain, not the largest number in a table.
AI compute atlas
The machinery behind
every workload.
A researched directory of current accelerators, cloud silicon, servers and rack-scale systems for AI inference, model training, reasoning and high-performance computing.
Compute architecture
A workload is a
whole system.
Peak arithmetic is only one part of useful performance. Model shape, numerical format, memory capacity, bandwidth, interconnect, runtime maturity, power and operational access determine what the system can actually deliver.
SYSTEM MAP / 06 CONNECTED LAYERS
Start above the chip.
Finish beyond it.
- 06Service targetLatency, throughput, quality, reliability and cost
- 05Model and workloadTraining, decode, reasoning, retrieval, vision or simulation
- 04Framework and runtimeGraph compiler, kernels, serving and orchestration
- 03Accelerator and memoryNumerical formats, capacity, bandwidth and locality
- 02Node and fabricScale-up links, scale-out network, storage and host
- 01Facility and operationsPower, cooling, availability, observability and support
System balance
Performance
is a chain.
A strong accelerator can still wait on memory, communication, software or the facility around it. Compare the complete path and identify the first real bottleneck.
Fit weights, activations, cache and training state without forcing harmful movement.
Keep compute engines supplied during memory-bound attention, decode and vector work.
Understand how quickly accelerators communicate inside a node, rack and cluster.
Check framework coverage, kernel quality, serving maturity and operational tooling.
Confirm rack density, cooling method, facility headroom and energy constraints early.
Compute systems directory
Compare the system.
Not the slogan.
The order balances capability, real availability and ecosystem reach. It is a deployment-relevance index, not a speed leaderboard. Peak figures remain vendor theoretical specifications unless stated otherwise.
Try another class, availability, deployment path or search term.
Deployment comparison
Place the workload
where it belongs.
Choose up to three systems in the directory for a field-by-field comparison. Before selection, the matrix explains the operational differences between common deployment paths.
0 of 3 selected
Memory and fabric
Move less.
Use more.
Large AI workloads are data-movement systems. Useful throughput depends on where weights, activations and cache live, how often they move and whether the interconnect keeps parallel work synchronized.
- 01 / LOCALITY
- Keep the hottest data close to the compute engines that consume it.
- 02 / SCALE-UP
- Measure the fabric that lets several accelerators behave like one larger system.
- 03 / SCALE-OUT
- Plan communication, storage and failure domains when work crosses nodes and racks.
Research method
Useful claims need
clear boundaries.
Every entry separates product form, availability, access model and manufacturer theory. Cross-vendor figures remain context, never a substitute for workload testing.
Define the unit
A chip, card, server, rack and cloud slice are different systems. Their totals must not be mixed.
Label precision
FP4, FP8, BF16, sparse and dense peaks describe different numerical conditions.
Verify access
Shipping hardware, cloud capacity, preview SDKs and announced plans receive different status labels.
Prefer sources
Specifications link to the vendor or standards body so teams can confirm current detail.
Test the workload
Model shape, batch, context, kernels and service goals decide sustained performance.
Use MLPerf precisely
Independent results need their exact model, scenario, accuracy, software and system configuration.
Independent evidence layer: MLPerf Inference v6.0 results ↗
Before selecting compute
Six questions before
the purchase order.
What is the service target?
Set latency, throughput, quality, availability and cost limits before choosing silicon.
What must fit in memory?
Account for weights, activations, key-value cache, optimizer state and safety margin.
How will the work scale?
Map tensor, pipeline, expert and data parallelism onto the real fabric topology.
Is the software ready?
Validate the model, precision, kernels, compiler, serving stack and observability path.
Can operations support it?
Confirm power, cooling, networking, storage, staffing and vendor support requirements.
Can the claim be reproduced?
Run representative tests with production prompts, batches, context and quality checks.
Engineering the full path
From compute choice
to working system.
ADOR.IS designs AI-enabled products, model-serving infrastructure, data systems, graphics platforms and operational software around measurable constraints.
[email protected] ↗