Loading the ridge

We weighed fifteen ways to read the same satellite tile.

Five architectures, three numeric precisions, one training recipe, ten classes of land cover. Then we measured what each answer costs in joules — the number model cards never print.

The problem

Training is a bill you pay once.

Inference is a subscription. A deployed land-cover model runs until it is replaced, on someone else's CPU, long after the paper. Almost nobody measures that, so nobody optimises it.

Why it needed measuring

Accuracy had stopped telling them apart.

Under one recipe and no per-model tuning, all five finish within a hair of each other. On a dataset this saturated, choosing by accuracy is choosing by noise.

0
What it changed

Same answer.
0 the electricity.

Three joules per thousand tiles against sixty-one. The deployment decision was never about accuracy — it was about the number nobody had bothered to publish.

The drive

Green AI is a measurement problem
before it is an engineering one.

Nothing here needed a new architecture. It needed weighing the ones we already have, on the hardware they actually run on, and publishing the result whole.

0 Energy spread at equal accuracy
0 Accuracy given up
0 Timed measurement windows
0 Tests over the artefacts

Same answer.
19× the electricity.

Five architectures, one training recipe, ten classes of land cover seen from orbit. Accuracy separates them by 1.25 points. Energy separates them by nineteen times.

Run a model

A model is trained once, then run forever.

Training is a bill you pay one time. Inference is a subscription, and almost no model card prints the price.

Accuracy had stopped telling the models apart.

Under one identical recipe, five architectures land within 1.25 points of each other. Choosing by accuracy is choosing by noise.

So we measured the other axis.

Five seeds each, three precisions, one machine, CPU runtime only, power sampled off the die during every timed window.

Same accuracy, nineteen times the energy.

3.18 joules per thousand tiles against 60.84. The cheap model gives up two thirds of one accuracy point to do it.

Fifteen configurations. One answer. Nineteen times the bill.

Five architectures, three numeric precisions. Put the cheapest and the dearest in front of the same thousand tiles and they differ on about seven of them.

Measured on one Apple M2 EfficientNet-Lite0static int8 ResNet-50fp32 baseline Difference
Energy per 1,000 tiles 3.18 J60.84 J 19× less
Top-1 accuracy 97.45%98.12% 0.67 pp worse
Latency, 95th percentile 0.43 ms12.24 ms 29× faster
Model on disk 3.8 MB94.0 MB 25× smaller

Accuracy is a mean over five seeds. Energy is sampled from on-die telemetry during each timed window.

Nobody prints the joules.

Model cards report accuracy to two decimal places and say nothing about power. The field buys tenths of a point with electricity it never counts.

Inference is the part that repeats

Training happens once. A deployed land-cover model runs until it is replaced. At a million tiles a month, the two configurations above differ by 7.7 grams of CO₂e per million — 0.42 against 8.13, on a world-average grid of 481 gCO₂e/kWh.

CPU is where most of it happens

Earth-observation pipelines run on whatever is available, which is usually a CPU without a GPU beside it. Every figure here is measured through the ONNX CPU runtime on one Apple M2, single-threaded.

The tiles stay 64 pixels

EuroSAT is natively 64×64. Upscaling to the backbones' 224×224 pretrain size costs roughly twelve times the CPU work and adds no information — and CPU work is the quantity under test.

One recipe, no favourites

Same optimiser, same twenty epochs, same augmentation, same learning rate for all five architectures. If a model does badly under a fixed budget, that is a finding about the model, not a reason to tune it until it wins.

Four results worth the electricity.

Every figure is read at page load from the committed benchmark. None was typed into this page.

19×

The gap that should decide the deployment

EfficientNet-Lite0 in static int8 costs 0.67 points against the ResNet-50 baseline and runs on a nineteenth of the energy, at 29 times lower p95 latency.

-64.66 pp

Quantisation is not a free lunch

The same post-training recipe is free on one architecture and destroys another. Watch it happen in the demo.

Dynamic int8 loses on both axes at once

Worse accuracy and worse energy than plain fp32, for every model measured. A negative result, published because it saves the next person the experiment.

Energy is mostly latency — but not only

The two track each other at r = 0.983. The residual runs against intuition: quantised models finish sooner while drawing more power per second than the models they replace.

Twenty tiles from the held-out test fold — the exact images the demo classifies.

Don't take the numbers. Run the models.

The demo downloads the benchmark's own ONNX graphs and runs them in your browser, on your machine, on tiles the models never trained on. Nothing is uploaded. Watch MobileNetV3-Small get it wrong — faster than anything else on the page.

Open the demo How it was measured