AI Chess

What A Chess Engine Actually Does When You Press Analyse

There is a moment, right after a loss, when you paste the PGN into Lichess and hit the analyse button, and a number appears: +1.62. It feels like a verdict. It feels like the position has been weighed on some cosmic scale and found wanting, and that your move 23 was Wrong in the way that 2+2=5 is wrong.

It isn’t that. Not even slightly.

Stockfish is not an oracle. It is a decision procedure: a program that generates candidate futures, scores them with a learned function, throws away the ones it judges unpromising, and hands you back an ordered list. The number on screen is the score of the best line it managed to find in the time you gave it. Change the time, change the number. Change the hardware, change the number. Ask it for three moves instead of one and you will often find that its second choice is 0.04 pawns behind its first, which is to say the two moves are indistinguishable and the engine has expressed a preference, not a truth.

Everything else you will ever learn about engine use depends on internalising this. So let’s take the machine apart.

The two halves

Ask how does a chess engine work and you get one honest answer with two parts: a search and an evaluation function. Search builds a tree of possible continuations. Evaluation puts a number on the leaves of that tree. Search then propagates those numbers back up using minimax, assuming both sides play the best line they can see, and the root move with the best backed-up score wins the vote.

That’s it. That’s the whole architecture. Everything sophisticated in Stockfish 17 is an optimisation of one of those two jobs.

The problem is scale. Chess has an average branching factor around 35, so a full tree ten plies deep (five moves each) is roughly 35^10, about 2.7 quadrillion positions. Stockfish on a decent laptop does maybe 2 to 5 million nodes per second per thread. Do the arithmetic and a brute-force ten-ply search takes years. Depth 30, which Lichess cloud analysis reaches routinely, is not remotely possible by brute force.

The engine survives by refusing to look

Alpha-beta pruning is the first trick. If you’ve found a line that gives you +0.3, and you start examining a second move where the opponent has a reply that drops you to -2.0, you don’t need to see the opponent’s other replies. One refutation is enough to discard the move. That alone cuts the effective branching factor from 35 to roughly 6 with good move ordering.

Stockfish goes much further, and this is the part club players never hear about. Null move pruning: pretend you pass your turn, and if the position is still winning for you, assume the real move is at least that good and cut the search short. Late move reductions: moves ordered late in the list get searched to shallower depth, on the theory that the move ordering is probably right. Futility pruning, razoring, singular extensions, aspiration windows that start the search assuming the score will land near the previous iteration’s result.

The net effect is an effective branching factor near 2. The engine is discarding more than 99% of legal continuations at every node, on the basis of heuristics that are usually right. Usually. Each of those pruning rules is a bet, tuned over millions of self-play games to be profitable on average. An engine reaching depth 30 has not examined all lines to 30 ply. It has examined a narrow, aggressively pruned sliver, with a seldepth (selective depth) often 10 plies deeper down forcing sequences and correspondingly shallower everywhere else.

The number itself is a guess

Since Stockfish 12 (September 2020), the evaluation function has been a neural network: NNUE, an efficiently-updatable architecture that reads the board through a king-relative input encoding and outputs a score in centipawns. The hand-written evaluation terms, bishop pair bonus, doubled pawn penalty, king safety tables, are gone. Stockfish 16 removed the classical evaluation entirely.

What the network learned is not “the truth about this position”. It learned to predict the outcome of a search from millions of labelled positions. A centipawn is not a physical quantity. Since Stockfish 14 you can run setoption name UCI_ShowWDL value true and get the honest version: win/draw/loss expectations in permille, calibrated at move 50 in a game between engines of Stockfish’s strength. A “+1.00” is roughly the point at which a strong engine converts often enough to matter. Between two 1400s it means approximately nothing.

Read an actual line

Fire up Stockfish from the command line and type uci, then position startpos, then go depth 24. You get a stream, and the last line looks like this:

info depth 24 seldepth 33 multipv 1 score cp 34 nodes 14221904 nps 2103244 
hashfull 512 tbhits 0 time 6762 pv e2e4 e7e5 g1f3 b8c6 f1b5 a7a6 f1e1
bestmove e2e4 ponder e7e5

Every field there is a claim about process, not about chess. depth 24 is nominal depth. seldepth 33 means some line went nine plies deeper. nodes 14221904 over time 6762 gives you the 2.1 million nps. hashfull 512 means the transposition table is 51.2% full, and if it hits 1000 the engine starts evicting useful work. pv is the principal variation, the line the engine believes both sides will play. The pv is the single most valuable thing on that line and the thing club players ignore hardest.

Now do the thing that actually teaches you something. Set setoption name MultiPV value 3 and re-run:

info depth 24 multipv 1 score cp 34 pv e2e4 e7e5 g1f3
info depth 24 multipv 2 score cp 31 pv d2d4 g8f6 c2c4
info depth 24 multipv 3 score cp 28 pv g1f3 d7d5 d2d4

Six centipawns separates 1.e4 from 1.Nf3. The engine has ranked them. It has not judged them. If you took the top line as gospel you would conclude that 1.e4 is objectively best and 1.Nf3 is an error, which is obviously absurd, and the absurdity is visible only because you asked for three lines instead of one.

The same pattern shows up in real openings. Put the Najdorf on the board after 1.e4 c5 2.Nf3 d6 3.d4 cxd4 4.Nxd4 Nf6 5.Nc3 a6 and run MultiPV 4. You will find 6.Bg5, 6.Be3, 6.h3 and 6.f3 clustered inside about a fifth of a pawn of each other, in whatever order that particular net and that particular depth happens to produce. The clustering is the information. The ordering is noise.

Where the procedure visibly breaks

Two failure modes are worth hunting deliberately, because seeing them once cures the oracle instinct permanently.

The horizon effect. Set up a position where you’re losing a piece in six moves and give the engine only depth 12. It will find a check, or a pawn push that forces a recapture, anything to shove the loss past the last node it looks at. Score reads -0.4. Run the same position to depth 28 and it reads -3.1. Nothing about the position changed. The engine simply stopped lying to itself once the truth fit inside the tree. Watch the eval as depth climbs: stability across depths 20 through 30 is evidence, a number that lurches at depth 26 is a warning.

Fortresses. Take an opposite-coloured-bishop ending where White is two pawns up but the pawns are blocked and Black’s bishop covers every entry square. Stockfish will sit there at +1.9 for as long as you let it. Minimax has no concept of “no progress is possible”; it sees material, the net was trained on positions where two extra pawns usually wins, and no search depth you can afford will reach the 60-move repetition that proves the draw. Syzygy tablebases solve this perfectly for seven pieces or fewer (18.4 TB of them, which Lichess queries for you), and not at all for eight.

Pattern-matching those two failure shapes is most of the skill. The rest of it, the calibration of what a given number means at a given depth in a given phase, is worth studying properly: the companion piece on reading engine output: evals, depth and lies walks through the calibration in detail.

Interrogation, not consultation

Consulting an engine means asking “what’s the eval” and accepting the answer. Interrogating it means running a protocol.

Here’s a workable one for reviewing your own game. Open the position where you felt uncomfortable, not the position where Lichess drew a red blunder icon (those are computed at shallow depth on a server budget and flag the symptom, not the cause). Set MultiPV to 3. Let it run to depth 28 or so, and watch whether the top three scores are spread or clustered. Then play out the principal variation on the board, move by move, and stop at the first move you don’t understand. That move is your lesson. The eval was never the lesson.

One more setting worth changing tonight: setoption name Threads value 4 and setoption name Hash value 1024. Default Stockfish runs on one thread with 16 MB of hash, and in the Lichess browser engine you’re getting a WebAssembly build that’s slower still. Give it four threads and a gigabyte and you’ll reach depth 30 in the time it previously took to reach depth 24, which is roughly the difference between a number you can trust and one you can’t.

Go find a drawn rook ending you lost. Run it at depth 35 and see what the machine says.