About

We don't tell you what should work. We test what actually works.

Hard Numbers exists because most engineering content is vibes. We publish benchmarks of AI models, programming languages, and frameworks — with the methodology, the failure modes, and the raw numbers. If we got it wrong, the methodology section will show how.

Every claim is tagged

Observed, Measured, Documented, Inferred, Hypothesized, or Unknown. We mark which is which so you know what to trust.

Benchmarks include methodology

Hardware, model versions, batch sizes, prompts — all specified. You can reproduce our numbers or disagree with them on the merits.

Code is open

Every experiment has a public repository. If we got it wrong, the commit history shows how.

No invented citations

Every URL in every article is real. Every API parameter is real. Every version is real. We mark 'Unknown' when we don't know.

Failure cases are first-class

Every article has a 'Failure cases' or 'Limitations' section. We don't hide what doesn't work.

Get in touch

Spot a mistake? Found a better way? Open an issue on the relevant repo, or reach out on Twitter.

← Back to articles