How it was measured.

A benchmark is worth exactly as much as its provenance, so the provenance ships with it. Every claim on this site can be traced from here to a file in the repository.

One recipe, five architectures, five seeds.

Every architecture is trained under an identical recipe. There is deliberately no supported way to give one model a tuned recipe of its own, because a benchmark where each entrant brings its own training budget measures the budget, not the architecture.

Split

EuroSAT ships no official train and test partition, so every published EuroSAT accuracy is measured against folds the reader cannot inspect. This split is generated deterministically from one seed, verified byte-identical across runs and across PYTHONHASHSEED values, and referenced by hash from every result row. A run aborts if the file stops matching its own checksum.

Train

Five architectures, five seeds each, under the single hashed recipe. Training used the GPU where one was available, purely to make the matrix tractable. No reported figure depends on the training device: accuracy comes from the saved checkpoints and every timing is measured later on the CPU.

Export

Each checkpoint is exported to ONNX in fp32 and quantised to int8 both dynamically and statically, with static calibration drawn from the training fold only. All 25 fp32 exports reproduce their checkpoint's accuracy exactly, asserted as a test rather than observed once: latency is measured on the exported graph while accuracy is attributed to the model, so a drift would mean every row described two different models.

Measure

Inference runs through ONNX Runtime's CPU execution provider only. Apple's GPU and CoreML providers are excluded deliberately, not merely left unused: the question is what CPU-only hardware achieves. Thread counts are pinned in both the runtime session and the environment, because BLAS pools ignore the session setting and unpinned threads are the most common reason CPU latency fails to reproduce.

Integrate

Energy is integrated from on-die package power sampled at 200 ms across each timed window. It is estimated, not metered at the wall, and this site says so everywhere it appears.

What the benchmark refuses to claim.

The limits are part of the result. Reading them is how you know which of these numbers travel to your hardware and which do not.

Excluded windows, and the confound.

Four of 300 measurement windows were excluded by criteria registered before any measurement was taken. None falls in the primary reporting configuration, so no headline figure changes when they are removed. The rows remain in the published data.

Benchmark hardware, not your machine

Read further

PROTOCOL.md records the measurement protocol, its outcomes and every deviation from the original pre-registration, including two criteria that were never instrumented.

DATASHEET.md documents the split and results artefacts following Datasheets for Datasets.

The repository carries the raw measurement rows, so every figure here can be recomputed.

Why any of this is checkable

The split, the recipe and every measurement window are committed to the repository, and the figures on this site are read from them at page load rather than typed into the markup. Eighty-two tests run over the artefacts on every change; several of them exist purely to fail if a number here stops matching the benchmark behind it.