Created by Ben O’Leary Co‑Founder & Chief Quantum Officer, Nirmata Holdings
603‑930‑2131 ben@nirmataholdings.com www.nirmataholdings.com
ATOM · V4 release · formula contract IY-v2.1 · source-grounded
Useful intelligence
per joule.
Not one magic number — a vector. Completed tasks per joule, net dollars per joule, watts per active user, latency at p99, and a bounded index that can be vetoed. Every input carries an evidence tier and a live low/base/high interval.
On the phrase. “Intelligence per joule” is an orchestration objective, not an SI unit. Intelligence has no accepted scalar measure, so this app substitutes a declared, auditable proxy: completed tasks on a named evaluation suite, divided by measured joules inside a declared boundary. Read the substitution as a design choice, not a discovery.
Six rings = six desirabilities. Arc length is the desirability, ring thickness is its weight. Exact values are in the synchronized inspector below — nothing important is read off the picture.
Command
Scenarios are labelled by what they actually are. The default is a measured reference with every delivery-convergence assumption switched off.
Bounded index
IY = 100·exp(Σ wⱼ·ln dⱼ) — a weighted geometric mean in log space, so a near-zero factor cannot be bought back by a strong one. Weights sum to exactly 1. The index is dimensionless and always relative to the frozen reference contract named above; it is never a physical rate.
Physical headline rates
MEASURED referenceYenergy and Yvalue are per completed task, not per attempt — attempts per completed task (na) is an explicit input, so retries cannot be hidden. Yvalue is allowed to go negative; loss-making configurations are shown as losses.
V4 release · formula contract IY-v2.1 · REF-2026.1
The methodology kernel
One objective. Three mathematical layers.
Every layer below is printed in full, substituted with the values currently loaded, and labelled with what it is allowed to be used for. Only layer C is the authoritative score.
On the two version numbers. V4 is the product release you are looking at. IY-v2.1 is the mathematical formula contract it runs on, frozen under reference REF-2026.1. V4 added this kernel, the baseline observatory and the reference contract; it did not change the index formula, the weights, the frozen targets or the reference basket, so every V3 index value still holds.
Symbol legend — every term, its unit and its role
| Symbol | Meaning | Unit | Current | Role / source |
|---|
Baseline observatory
The sourced numbers this model is calibrated against. Every row carries an evidence tier, its workload and boundary, an as-of date and a live link. Full detail, including the cross-source normalisation limits, is in Evidence & Method.
Per active user and whole-system vector
Every figure names its unit and its allocation basis. “Per active user” means divided by the declared active-user count — an accounting allocation, not a measurement of one person's hardware.
Executive controls
Baseline energy, usage, task success, cache hit rate and serving gain move almost all of the vector. Everything else is available below, and in full in the Measurement Lab.
Advanced controls latency p99 · throughput @ concurrency · tokens · split prices · privacy · routing · network · attempts/task · rebound
Delivery convergence
all default to 0Fleet totals
Measurement Lab
Direct measured energy is the default and is bound to a named workload and boundary. The bottom-up model is optional and is reconciled, never silently mixed.
Energy accounting mode
Workload & task success
Accuracy and task-fit are merged into one measured end-to-end task success rate Q on a declared suite. Rating the same capability twice was a double-count.
Latency & throughput
TTFT p99, TPOT p99, system throughput at a stated concurrency and SLO attainment are four different things and are kept apart.
Cost & carbon
list pricesInput, output and cached-input tokens price separately. A cache hit cuts both energy and token cost — the 2026 app cut neither correctly.
Privacy & compliance
Compliance is a gate, not a score you can trade away. Fail any required control and the index is vetoed to zero with the failing control named.
Uncertainty (always on)
Sensitivity ranking
Index weights
stakeholder choiceWeights are a governance decision, not a measurement. They are renormalised to sum to 1 and shown as a normalised percentage next to each slider.
Equation Trace
Every headline number, substituted with the numbers actually in play and carried at full precision.
Phrase → physics
The blueprint asked for “intelligence per joule” and “tokens per watt per user”. Neither is an SI quantity as written. This is the exact path from the sentence to something you can measure and defend.
Equation inspector
Log-contribution waterfall
Each bar is wⱼ·ln dⱼ. They sum to ln(IY/100) by construction — the identity is re-checked on every recompute and shown above.
Dimension & unit contract
| Quantity | Unit | Derivation |
|---|
Unit invariance. Every dⱼ is a ratio of two quantities in the same unit, so the index is dimensionless and unchanged if you express energy in joules instead of watt-hours, or money in cents instead of dollars. Physical rates keep their units and convert exactly. The validation harness proves both (tests DIM-1 and DIM-2).
Landauer guard
physics floorLegacy / literal interpretation
comparison only — not authoritativeBoth formulas below are kept so the correction is auditable. Neither feeds the index, the vector or any provider row.
Retired V1 — V = (A·S·P)/(C·E)
Literal blueprint form — IYI = (A·Sₙ·P·Tu)/(Cₙ·Eu)
Provider Arena
No single score should hide a tradeoff. Provider rows are scenario assumptions, evidence-linked on capability and explicitly modelled on outcome.
Pareto frontier
Delivery-path comparison
| Delivery path | Baseline Whper attempt | Adjusted Whper attempt | Wh / userper day | W / user | Φ tokens/J | Tasks / J | Index | Evidence | Pareto | Actions |
|---|
Every row is modelled: the same calculate() runs on your current
configuration with that path's deltas applied. “Evidence” is the share of that path's deltas carrying a real
citation; the rest print as n.a. rather than receive an invented source. Edit any row and the table,
the frontier and the leader all move.
Why Akamai can win
capability → lever → evidence → pilotDeliberately not “Akamai wins”. This is the falsifiable version: each capability maps to the model lever it would move, the evidence tier that supports it today, and the one measurement a pilot must produce to promote that row out of inference.
| Capability | Model lever | Evidence today | Required pilot measurement |
|---|
Evidence & Method
What is measured, what is sourced, what is modelled, and what is simply not known.
Reference contract — REF-2026.1
An index without a frozen, inspectable reference is a system scoring itself. This is the whole contract: what is fixed, when it took effect, when it is reviewed, and what counts as a change to it.
Methodology changelog
| Version | Date | Change |
|---|
Baseline observatory — full detail
Transcribed from the fetched-source baseline brief compiled 2026-08-06. Where a fetched page did not state a value it is printed n.a. rather than filled in.
Cross-source normalisation limits
Energy, throughput and price are three different axes. They are not mutually convertible and this app does not imply that they are.
Evidence tiers
An aggregate scenario is never labelled sourced if any decisive input inside it is modelled or conceptual. The scenario badge on the Command tab reports the weakest tier it depends on.
Source ledger
| Input | Value | Tier | Source |
|---|
Method
Evidence gaps and honest limits
n.a.- No public per-token energy measurement exists for the frontier proprietary models. The 0.24 Wh median comes from Google's own Gemini production measurement and is not transferable to another vendor's stack without re-measurement. Google, Measuring the environmental impact of delivering AI.
- Edge/CDN inference energy is not published by any vendor at task granularity. Every delivery-convergence input in this app is therefore CONCEPTUAL, defaults to zero, and is never presented as a measurement.
- MLPerf Inference power results are optional and sparsely submitted, so latency and energy usually come from different runs on different systems. Mixing them is exactly the defect this rebuild refuses. MLCommons datacenter inference.
- Task success rate Q is only as good as the declared suite. A suite name is required; a number without one is meaningless. HELM makes the multi-metric case.
- Embodied carbon and water are out of scope in every boundary offered here, and the app says so rather than quietly excluding them.
- The Monte Carlo interval is a coverage interval from declared triangular priors, not a confidence interval and not a posterior. It inherits every bias in those priors. JCGM 101 (GUM Supplement 1).
- The frontier scenarios are not benchmarks. No vendor in the comparison has published task-level joules for their edge product; the deltas are modelled by us.