What it costs, and what to deploy.
Quantisation is close to free on one architecture and destroys another. Below that, set the constraints a real deployment has and read the cheapest configuration that satisfies them.
Quantising is not a safe default.
Paired per-seed change against each model's own fp32 export, with a Student-t interval over seeds.
Every measured precision, not just the two above. An interval spanning zero means no detectable change.
Set the constraints. Read the frontier.
Give the deployment its real limits and the cheapest configuration that satisfies them is the answer. Where nothing qualifies, that is also an answer, and this page says so rather than relaxing the question.
The dashed line is the Pareto frontier: nothing measured beats a point on it for both energy and accuracy. Amber marks the configurations meeting your constraints. The horizontal axis is logarithmic.
What the frontier is for
Every point above is a configuration that was actually measured. The line joins the ones nothing else beats on both axes at once — pick a constraint and the cheapest configuration that still satisfies it is the one to deploy. There is no interpolation here and no model that was not run.