How AI Benchmarks Work and Why You Should Read Them Sceptically
AI benchmarks can reveal genuine strengths, but understanding their tasks, scoring rules and hidden assumptions is essential before trusting a leaderboard.
Archive
Everything published here, newest first. Drafts are private until a person has verified their facts, sources and attribution, so this list only ever shows reviewed work.
25 articles · page 3 of 3
AI benchmarks can reveal genuine strengths, but understanding their tasks, scoring rules and hidden assumptions is essential before trusting a leaderboard.
Small language models can be the better choice when a task is narrow, resources are limited and reliability comes from a well-designed workflow.
Chain-of-thought prompting can help AI tackle multi-step problems, but its real value comes from useful decomposition, visible checks and knowing when a simpler prompt is better.
Tree of Thoughts turns difficult AI tasks into a controlled search through alternatives, with explicit checks, pruning and stopping rules.
Prompt chaining turns a large AI request into smaller, checkable stages, making complex work easier to inspect, correct and repeat.
Learn how large language models turn text into tokens, build numerical representations and use attention to generate answers, with worked examples and practical exercises.
How Promptbreeder and APE use evolutionary algorithms to mutate and self-improve LLM prompts.