Why Stockfish’s Best Move Is Often Useless Advice To You
You lose a game, you paste it into the Lichess analysis board, and a big blue arrow appears. Nxf7! Evaluation: +1.2. You nod, you close the tab, and you have learned precisely nothing you can use in the next round of your club championship.
This happens because Stockfish is answering a question you did not ask. You asked: what should I have played? It answered: what is the move that holds up against a perfect opponent who will find every defensive resource, including the one on move 14 that involves stepping the king into a corner it appears to have no business visiting. Those are different questions with different answers, and the gap between them is where most club-level improvement gets thrown away.
The engine’s top line is a proof, not a plan
Stockfish 17 does not play chess against you. It plays chess against a copy of itself, and then reports the outcome of that imaginary game. Every eval you see is the endpoint of a variation in which both sides defend at roughly 3600 strength.
That has a specific, predictable consequence: the first line is biased towards sharp continuations whose soundness rests on a defensive resource buried deep in the tree. The engine is happy to walk into a position that is only tenable because of a zwischenzug at ply 15, because for the engine that zwischenzug is not hard. It is a node. You, meanwhile, need to find it under a clock, after 90 minutes, holding a shape in your head with six pieces still on the board.
Here is the asymmetry that nobody explains to club players. Attacking resources are pattern-shaped and often findable: sacrifices on h7, rook lifts, back-rank tricks. Defensive resources that far down the tree are usually anti-pattern moves: a king walk, an interposition that loses a tempo, a counter-sac that returns material at the worst-looking moment. Those are the moves human beings do not see. So when Stockfish says a line is fine for the defender, it is telling you something true about chess and something false about your Tuesday night.
Worked example: the Fried Liver, from the wrong end
Take the most over-analysed position in club chess.
1. e4 e5 2. Nf3 Nc6 3. Bc4 Nf6 4. Ng5 d5 5. exd5
You are Black. Two candidate moves: the natural recapture 5...Nxd5, and the ugly-looking 5...Na5, which sends a developed knight to the rim to trade off White’s bishop.
Run this on the Lichess board with local Stockfish at default settings and watch what happens as depth climbs. Around depth 16 to 20 the two moves sit close together, something like +0.5 against +0.8, close enough that the engine’s ordering can flip between iterations. Keep it running past depth 35 with four threads and the picture sharpens: 5...Nxd5 6.Nxf7 Kxf7 7.Qf3+ Ke6 8.Nc3 Nb4 and Black is, objectively, holding a piece for two pawns and a shredded king. The evaluation stays inside a pawn.
Now ask yourself the only question that matters. How many times in your life have you sat on the black side of 8...Nb4 9.Qe4 c6 10.a3 Na6 11.d4 Nc7 12.Bf4 and defended it correctly? The eval is honest. The advice is worthless. Below about 2000, Black’s practical score in that structure is a bloodbath, because the defence is a sequence of only-moves and the attack is a sequence of natural ones.
5...Na5 is the move you should study. It is not the engine’s darling, it usually shows as the second or third line, and it hands White a small stable edge instead of a chaotic near-equality. It is also the move you can learn in twenty minutes and play correctly for the next decade.
That is the whole shape of the problem: the engine optimises for the evaluation at the end of the line, and you need to optimise for the evaluation at the end of the game.
Worked example: the endgame where 0.00 is a lie about you
Set up a textbook Philidor position. Black king on f8, black rook on a6, white king on e6, white rook on h1, white pawn on e5, Black to move.
Stockfish’s output is brutal in its clarity:
1...Ra1→ around −4.5 for Black, dead lost1...Ke8→ around −4.5, also lost1...Ra7,1...Rb6,1...Rc6→0.00
The engine has told you the truth and taught you nothing. The reason the rook must sit on the sixth rank is that it stops the white king reaching the seventh, and the reason that matters only becomes visible when the pawn advances to e6 and the rook swings to a1 to check from behind. That is the actual lesson, and it lives about eleven plies past the move the engine showed you. A 0.00 with no name attached to it is a fact you cannot store, retrieve, or apply.
How to actually interrogate the thing
Stop reading one arrow. Start reading distributions.
Turn MultiPV up. On Lichess, the gear icon in the analysis panel has a “Multiple lines” slider that goes to 5. On Chess.com’s analysis board, raise the engine line count. If you run a UCI engine directly (Nibbler is excellent for this, as is en-croissant), you are typing:
setoption name Threads value 6
setoption name MultiPV value 5
setoption name UCI_ShowWDL value true
go depth 32
And reading output that looks like this:
info depth 32 multipv 1 score cp 74 wdl 168 806 26 pv f3e5 c6e5 d4e5
info depth 32 multipv 2 score cp 61 wdl 141 830 29 pv c1e3 f8e8 a2a4
info depth 32 multipv 3 score cp 58 wdl 133 838 29 pv d1c2 g7g6 b2b4
Three moves inside 0.16 of a pawn, with win probabilities inside three percentage points. There is no “best move” here. There is a set of good moves, and you get to pick on human grounds: which one leads to a structure you understand, which one gives your opponent the most chances to go wrong, which one you can play quickly.
Read the WDL numbers, not just the centipawns. wdl 168 806 26 means 168 wins, 806 draws, 26 losses per thousand at that evaluation. A +0.74 that is 80% draws is a completely different practical object from a +0.74 that is 40% decisive. If the normalisation behind those numbers is new to you, the pillar piece on reading engine output, evals, depth and lies covers how Stockfish’s +1.00 got pegged to a 50% win expectancy and why older engine numbers do not translate.
Watch the line at multiple depths, not one. Run depth 18, then 25, then 35, and note which moves move. A recommendation that is stable from depth 18 upwards is a move rooted in something structural you can name. A recommendation that only appears at depth 30 and vanishes if you nudge the position is resting on a tactical resource you would never have found. That instability is the single most useful diagnostic signal the engine gives you, and almost nobody looks at it.
Handicap the engine deliberately. Stockfish exposes UCI_LimitStrength and UCI_Elo, with the Elo range starting at 1320. Set it near your own rating, play the position out, and you get a far better model of what will actually happen to you than any depth-40 eval. Better still, install Maia through Nibbler. Maia’s networks (maia1100, maia1500, maia1900) are trained to predict the move a human at that rating band plays, not the best move. Put Stockfish and maia1500 side by side on your losing position and the question changes from “what was best” to “what was I likely to do, and what would my opponent likely have done”, which is a question with an answer you can train against.
Turning this into a plan for the week
Keep a notebook of positions, not moves. For each game you analyse, find the moment the evaluation swung by more than 1.0 and write down three things: the top engine line, the second line, and one sentence on why the second line is the one you would trust yourself to execute. If you cannot write that sentence, you have not understood the position, and the engine’s arrow was decoration.
Set a gap threshold and hold yourself to it. Under 0.30 between line 1 and line 2, treat them as equal and choose on style. Between 0.30 and 0.80, ask whether the top line’s advantage survives a 1500-rated defence (the Maia test, the UCI_Elo test). Above 0.80, the engine has probably found something real and you should work out the tactical idea by hand before you look at the variation.
Then go and play the second-choice move in twenty blitz games and see what happens to your results. My bet is that the arrow you have been ignoring wins you more points this season than the one you have been worshipping.