Leela Versus Stockfish: Two Different Kinds Of Wrong
Here is a claim that will sound heretical if you have spent the last few years treating engine output as revealed truth: both Stockfish and Leela Chess Zero are routinely, predictably wrong about positions, and they are wrong in opposite directions. Stockfish will hand you a +0.00 on a position where you are strategically doomed. Leela will hand you a 0.55 win probability on a position where you get mated in nine. Neither is malfunctioning. Each is behaving exactly as its search architecture demands.
Learning to spot which failure mode you are looking at is the single highest-leverage engine skill a club player can pick up, and it costs you nothing but the habit of opening a second analysis tab.
The two machines, in one paragraph each
Stockfish searches with alpha-beta and a huge amount of pruning. It walks a tree of variations, scores leaf nodes with a small neural network (NNUE) that evaluates a static position in microseconds, and aggressively discards branches it can prove are worse than something it already has. Because it evaluates millions of nodes per second and because the pruning rules are tuned for concrete material and king safety, it is a forcing-line detection machine. Give it a position with a perpetual check, an unstoppable mating net, a knight fork eleven plies deep, and it will find it while you are still reaching for your coffee.
Leela does something else entirely. A large policy network looks at a position and proposes candidate moves with probabilities: “this is a Nf5 kind of position, 34%; also consider h4, 11%.” A value head guesses the win/draw/loss chances from that position without searching at all. Monte Carlo tree search then expands the promising branches, backing up averaged results rather than a single minimax value. Leela at 40,000 nodes has seen fewer positions than Stockfish sees in a tenth of a second, but every one of those positions was chosen by a network trained on tens of millions of self-play games where long-term compensation actually paid off.
If you want the full taxonomy of engine families, including where Maia and bare NNUE fit, the pillar on engines compared covers it. This piece is narrower: what each engine gets wrong, and how you exploit that.
Failure mode one: Leela and the forced line it never sampled
MCTS has a structural blind spot. A move only gets explored if the policy network gives it enough prior probability to be worth a visit, or if the search runs long enough for the exploration term to force a look. Sacrifices that work for exactly one concrete reason, with a single narrow follow-up, are precisely the moves policy networks underweight. The policy head learned patterns. A one-off tactical shot is not a pattern.
The classic demonstration is the Réti-style fortress and its cousins, but the more useful club-level version is any position with a long forced sequence and a single saving resource. Take a rook-and-pawn endgame where the defence is a stalemate trick eight moves deep. Run it and watch:
Leela (BT4, 40k nodes, WDL output)
eval -2.31 W 4.1% D 27.6% L 68.3%
best Kg7 (policy prior 41%)
Stockfish 17, depth 48
eval 0.00 tbhits 12,410,772
best Rh2+ with forced repetition / stalemate resource
Leela’s number is not a rounding error, it is a different answer. Its policy head never gave the saving move enough prior probability to receive visits, so those visits went to the natural-looking king move, and the value head dutifully reported that the natural move loses. Correct assessment of the wrong move.
You will meet this constantly in three places. Sharp forcing tactics in the middlegame, especially quiet in-between moves that only work because of one geometric detail. Endgames with fewer than seven pieces, where Stockfish’s Syzygy tablebase hits give it literal perfect play and Leela’s network is guessing. And any position where the correct move looks absurd: an underpromotion, a king walk into the open, a queen sacrifice for a positional bind that resolves in a specific mate.
Practical rule: if the position is forcing, trust Stockfish. Captures available, checks available, material imbalance in flux, fewer than seven pieces on the board. Stockfish is not guessing there, it is calculating, and its tablebase access makes late endgames non-negotiable.
Failure mode two: Stockfish and the compensation it prunes away
Now flip it. Consider the kind of position club players throw away every week: you have sacrificed a pawn for a bind, the opponent’s pieces are passive, and there is no immediate breakthrough. Twenty moves of slow squeezing might win it. Alpha-beta search struggles here for a structural reason, not a tuning reason. To see the payoff of a bind, you need to search past the horizon where the bind converts. Stockfish prunes the quiet moves that constitute the squeeze because, at each individual ply, they look like they do nothing.
A typical King’s Indian or Benoni structure where Black has given up material for dark-square control shows the split:
Position: Black a pawn down, knight on e5, bishop on g7,
White's pieces tied to defending d3 and b2
Stockfish 17, depth 44, 1 thread, 512MB hash
eval +0.72 (White better, extra pawn)
Leela BT4, 25k nodes
eval -0.18 W 28.4% D 49.1% L 22.5%
(Black's practical chances better than the pawn count)
Which one is right? At depth 60 with 32 threads and an hour, Stockfish often drifts toward Leela’s assessment. That drift is the tell. When Stockfish’s eval moves significantly as depth increases past 30, you are looking at a position where its static understanding disagrees with its search, and Leela’s judgement is worth weighting. If the eval is rock stable from depth 22 to depth 45, the number is probably real.
The other Stockfish blind spot worth naming is the closed position. Stockfish’s pruning assumes moves matter. In a locked structure where the plan is a fifteen-move manoeuvre to reroute a knight from b1 to f5, it will show something close to level and shrug. Leela’s policy network, trained on self-play where those manoeuvres decided games, will often show a 60/35/5 WDL split favouring the side with the better plan. That WDL split is the information. A raw centipawn number throws it away.
How to actually interrogate both
Get Leela running. On Lichess analysis you have Stockfish 17 NNUE in the browser and that is it. To run both properly, install Nibbler or download the Leela binary plus a network (BT4 or a recent T3 net) and drive it from Arena, Cute Chess, or Nibbler’s GUI. Nibbler is worth the fifteen minutes because it shows policy priors as percentages on each move, which is the most underrated training signal available.
Set both engines up to answer questions, not to confirm your move. Concretely:
| What you are looking at | Run this | Read this |
|---|---|---|
| Tactical shot, forcing lines | Stockfish, depth 35+ | Centipawn eval, PV to the end |
| Pawn sac for a bind | Leela, 30k+ nodes | WDL split, not the cp conversion |
| Endgame, ≤7 pieces | Stockfish with Syzygy | Tablebase result, ignore everything else |
| Opening choice for you | Leela policy priors | Top 3 priors = human-plausible plans |
| “Am I lost or just worse?” | Both | If they disagree by 1.0+, dig in |
That last row is the one that changes your training. Disagreement between the engines is not noise to be averaged away, it is a flag that says this position contains something a human can misjudge. Those are exactly the positions worth 20 minutes of your own calculation before you look at either engine’s answer.
Two more concrete habits. First, use MultiPV 4 on Stockfish rather than reading only the top move. You want to know whether the best move is best by 0.05 or by 1.8, because that tells you whether the position demands precision or forgives you. Second, when Leela shows a WDL split like W 31% D 58% L 11%, read the draw percentage. A position that is 58% drawn is a position where your practical winning attempt has to accept risk. Centipawns hide that. Stockfish’s +0.34 could be that near-dead draw or it could be a position with 70% winning chances against a human, and the cp scale will not tell you which.
The reframe that makes both useful
The centipawn number is a compression of two different things: how likely you are to win, and how much material and structure you have. Stockfish’s number leans on the second. Leela’s WDL output reports the first directly, which is why it is more useful for deciding what to play against a human and less useful for deciding whether a combination works.
So stop asking “what does the engine say” and start asking which question you have. Is this a calculation problem or a judgement problem? If you cannot tell, that itself is diagnostic: run both, and let the size of their disagreement tell you which muscle you need to train. A club player who consistently gets Stockfish-shaped errors (missed forks, missed defensive resources) needs tactics volume. One who gets Leela-shaped errors (pawn grabs that hand over the initiative, trading into structurally lost endgames) needs a different diet entirely, and no amount of puzzle rush will fix it.
Next time you analyse a loss, find the move where the eval swung, then check whether the engines agreed about the position before the swing. If they did, you missed something concrete and calculable. If they did not, you walked into a position that even the machines find genuinely hard, and your opponent probably did not outplay you so much as get luckier about which kind of wrong they were.