Tags
chess, chess analytics, chess history, engine analysis, expected score, performance metrics, Stockfish 18, WDL evaluation, world championship
- CHESS ANALYTICS 00.2: Methods, Part 3: How to Read the Output Files
- CHESS ANALYTICS 00.0: List of Other Chess Analytics Articles
- 1. The Basic Units of the Analysis
- 2. Expected Score
- 3. Expected-Score Loss
- 4. WDL Accuracy
- 5. Game Accuracy
- 6. Mutual Accuracy
- 7. Performance Quality
- 8. Dominance
- 9. Volatility
- 10. RMS Expected-Score Loss
- 11. Error Concentration
- 12. Score, Expected Score, and Conversion
- 13. HardRAP and SoftRAP
- 14. How the Metrics Form One Whole
- 15. Why WDL Is Preferable Here to Pawn Evaluation
- 16. Why This Series Is Worth Doing
2026-06-12, added Part 2:
- CHESS ANALYTICS 00.1: Methods, Part 2 — Post-WDL Difficulty Metrics
- 1. WDL Metrics: Measuring Lost Scoring Chances
- 2. All-Legal Difficulty: How Narrow Was the Position?
- 3. Difficulty-Weighted Accuracy: Accuracy Under Pressure
- 4. CCP Difficulty: Checks, Captures, and Promotions
- 5. CCP-v2: Recursive Forcing-Line Search Difficulty
- 6. CCP-v2 Vision: Looking for Quiet Ideas Before Forcing Lines
- 7. CCP-v2 Two-Stage Probe/Confirm: Search Cheaply, Then Confirm Seriously
- 8. CCP-v2 Root-Branch Diagnostics: What Did the Search Actually Prefer?
- 9. CCP-v2 Log-Critical Moves: Where Difficulty Meets Damage
- 10. What the Post-WDL Metrics Add
- 11. Short Chronological Summary of Analyzer v3.6.1
- 12. Practical Reading Guide
This article series studies World-Championship matches and World-Championship qualification runs with a modern engine-based method.
The basic idea is simple:
Put every move of great historical matches under Stockfish 18, translate each position into win/draw/loss chances, and then ask: who preserved winning chances better, who lost chances more often, who created volatility, who converted chances into points, and who stayed more consistent?
Stockfish 18 is far stronger than any human player. Its strength is so far above human World Champions that comparing it to humans by normal Elo becomes difficult. This makes it useful as a reference point: not because it “understands chess like a human,” but because it gives a very strong, consistent measuring stick.
The purpose is not to reduce chess greatness to one number. The purpose is to create a performance profile:
Continue reading