Every actively-used benchmark from SaylorTwift/llm-benchmark-usage, rendered as a card colored by category. Rarity — Common, Uncommon, Rare, Rare Holo, Ultra Rare, Rainbow Rare, up to a one-of-a-kind Secret Rare — is set by how many major labs' latest release (OpenAI, Anthropic, Kimi, Z.ai, DeepSeek, Qwen, Meta) report that benchmark: the more labs, the stronger the foil and glow, and the higher it floats in the deck. The "Used by" row shows which labs those are, and a card links to a "🏆 Leaderboard" instead of a plain HF dataset when its dataset is officially registered on the Hub. Move your mouse over a card to catch the shine — tap/hold on mobile.
Showing benchmarks used by 3+ models, with at least one report in the last 6 months —