We don't tell you what should work. We test what actually works.
Hard Numbers exists because most engineering content is vibes. We publish benchmarks of AI models, programming languages, and frameworks — with the methodology, the failure modes, and the raw numbers. If we got it wrong, the methodology section will show how.
Every claim is tagged
Observed, Measured, Documented, Inferred, Hypothesized, or Unknown. We mark which is which so you know what to trust.
Benchmarks include methodology
Hardware, model versions, batch sizes, prompts — all specified. You can reproduce our numbers or disagree with them on the merits.
Code is open
Every experiment has a public repository. If we got it wrong, the commit history shows how.
No invented citations
Every URL in every article is real. Every API parameter is real. Every version is real. We mark 'Unknown' when we don't know.
Failure cases are first-class
Every article has a 'Failure cases' or 'Limitations' section. We don't hide what doesn't work.
Get in touch
Spot a mistake? Found a better way? Open an issue on the relevant repo, or reach out on Twitter.