A model is trained once, then run forever.
Training is a bill you pay one time. Inference is a subscription, and almost no model card prints the price.
Accuracy had stopped telling the models apart.
Under one identical recipe, five architectures land within 1.25 points of each other. Choosing by accuracy is choosing by noise.
So we measured the other axis.
Five seeds each, three precisions, one machine, CPU runtime only, power sampled off the die during every timed window.
Same accuracy, nineteen times the energy.
3.18 joules per thousand tiles against 60.84. The cheap model gives up two thirds of one accuracy point to do it.
Fifteen configurations. One answer. Nineteen times the bill.
Five architectures, three numeric precisions. Put the cheapest and the dearest in front of the same thousand tiles and they differ on about seven of them.
| Measured on one Apple M2 | EfficientNet-Lite0static int8 | ResNet-50fp32 baseline | Difference |
|---|---|---|---|
| Energy per 1,000 tiles | 3.18 J | 60.84 J | 19× less |
| Top-1 accuracy | 97.45% | 98.12% | 0.67 pp worse |
| Latency, 95th percentile | 0.43 ms | 12.24 ms | 29× faster |
| Model on disk | 3.8 MB | 94.0 MB | 25× smaller |
Accuracy is a mean over five seeds. Energy is sampled from on-die telemetry during each timed window.
Nobody prints the joules.
Model cards report accuracy to two decimal places and say nothing about power. The field buys tenths of a point with electricity it never counts.
Inference is the part that repeats
Training happens once. A deployed land-cover model runs until it is replaced. At a million tiles a month, the two configurations above differ by 7.7 grams of CO₂e per million — 0.42 against 8.13, on a world-average grid of 481 gCO₂e/kWh.
CPU is where most of it happens
Earth-observation pipelines run on whatever is available, which is usually a CPU without a GPU beside it. Every figure here is measured through the ONNX CPU runtime on one Apple M2, single-threaded.
The tiles stay 64 pixels
EuroSAT is natively 64×64. Upscaling to the backbones' 224×224 pretrain size costs roughly twelve times the CPU work and adds no information — and CPU work is the quantity under test.
One recipe, no favourites
Same optimiser, same twenty epochs, same augmentation, same learning rate for all five architectures. If a model does badly under a fixed budget, that is a finding about the model, not a reason to tune it until it wins.
Four results worth the electricity.
Every figure is read at page load from the committed benchmark. None was typed into this page.
The gap that should decide the deployment
EfficientNet-Lite0 in static int8 costs 0.67 points against the ResNet-50 baseline and runs on a nineteenth of the energy, at 29 times lower p95 latency.
Quantisation is not a free lunch
The same post-training recipe is free on one architecture and destroys another. Watch it happen in the demo.
Dynamic int8 loses on both axes at once
Worse accuracy and worse energy than plain fp32, for every model measured. A negative result, published because it saves the next person the experiment.
Energy is mostly latency — but not only
The two track each other at r = 0.983. The residual runs against intuition: quantised models finish sooner while drawing more power per second than the models they replace.
Twenty tiles from the held-out test fold — the exact images the demo classifies.
Don't take the numbers. Run the models.
The demo downloads the benchmark's own ONNX graphs and runs them in your browser, on your machine, on tiles the models never trained on. Nothing is uploaded. Watch MobileNetV3-Small get it wrong — faster than anything else on the page.