AI Chess

Maia Explained: The Engine Trained To Play Like You

Every engine you have ever used is trying to beat you. That is the whole design brief. Stockfish 17 running at depth 30 on your laptop plays a move that a grandmaster would need an hour to justify, and when you ask it why your move was bad, it answers in a language built for a different species. You played 14…Nd7 and lost 0.9 pawns of evaluation. Fine. But you didn’t play 14…Nd7 because you miscalculated a nine-move tactical sequence. You played it because the knight looked loose and you wanted it defended, and Stockfish has nothing to say about that, because Stockfish has never once wanted a knight to look defended.

Maia is different in a way that sounds small and isn’t. It was trained to predict what a human of a specific rating would play next. Not the best move. The likely move. That single change in objective turns an engine from a judge into a mirror.

What Maia actually optimises for

Maia came out of a 2020 research collaboration between the University of Toronto, Cornell and Microsoft Research (the paper is Aligning Superhuman AI with Human Behavior: Chess as a Model System, Kaggle/KDD 2020). The team took the Leela Chess Zero architecture, which is a residual neural network that Leela uses for its own superhuman play, and retrained it on something completely different: roughly 12 million human games from the Lichess database, bucketed by rating.

Nine separate networks came out of that process, named maia-1100 through maia-1900 in 100-point increments. Each one was trained only on games where both players sat within 50 rating points of that band. maia-1500 has seen a great many 1500-rated games and essentially nothing else.

The scoring function is the part that matters. Stockfish is optimised for centipawn advantage. Maia is optimised for move-matching accuracy: given a position drawn from a real human game, what percentage of the time does the network’s top choice equal the move the human actually played? On held-out test positions, the matching networks hit roughly 50-52% at their target rating. Stockfish, tuned down with a UCI_Elo limit to the same strength, manages around 33-36% on the same positions. Leela at low node counts does slightly better, mid-40s, but still loses to the network trained specifically to imitate.

Fifty percent doesn’t sound impressive until you sit with it. In an average middlegame position with 30 legal moves, Maia calls the exact human move in one out of every two positions, on a single forward pass with no search at all. And the accuracy curve peaks at the right place: maia-1100 predicts 1100-rated players better than maia-1900 does, and maia-1900 predicts 1900-rated players better than maia-1100 does. The networks are genuinely rating-specific rather than just being weaker versions of each other.

Why “one forward pass, no search” is the whole trick

Here is the thing most people miss when they first try Maia and conclude it feels weird. Maia runs at one node. It does not search. It looks at the position, runs it through the network, and outputs a probability distribution over legal moves. That’s it.

If you run Maia with a normal search tree of say 10,000 nodes, you break it. It gets stronger and simultaneously less human, because search finds tactics that no 1300-rated player would find, and the whole value proposition evaporates. In Lc0 terms your config wants:

lc0 --weights=maia-1500.pb.gz --nodes=1 --backend=eigen

Or in a GUI, set node limit to 1 and turn off any “ponder” or time-based thinking. If your Maia is thinking for two seconds a move, something is misconfigured.

The single-pass design is why Maia’s blunders are structured rather than random. A Stockfish handicapped to 1500 blunders by randomly sampling a worse move from its evaluated list. It plays ten strong moves then throws a rook for no reason, which is nothing like a real 1500 and teaches you nothing. Maia blunders the way humans blunder: it misses backwards moves, it walks into the same knight fork pattern repeatedly, it captures towards the centre when it shouldn’t, it gets seduced by pins.

A worked example: the move Maia predicts and you make

Take a position from the Italian Game, after 1.e4 e5 2.Nf3 Nc6 3.Bc4 Bc5 4.c3 Nf6 5.d4 exd4 6.cxd4 Bb4+ 7.Nc3 Nxe4 8.O-O Bxc3 9.d5. This is the Møller Attack. White has sacrificed a piece and the position is genuinely sharp.

Run the position through the different Maias and you get something like this distribution over Black’s ninth move:

Position: r1bqk2r/pppp1ppp/2n5/3P4/2B1n3/2b2N2/PP3PPP/R1BQ1RK1 b kq -

maia-1100 policy:      Bf6  31%   Ne5  19%   Bxa1 14%   Nd6  9%
maia-1500 policy:      Bf6  38%   Ne5  17%   Bxa1 11%   Qe7  7%
maia-1900 policy:      Bf6  44%   Ne5  12%   Nd6   9%   Bxa1 8%
Stockfish 17 (d=35):   Bf6  (the theoretical main line, eval ~0.0)

Every network gets the right answer at the top. But look at the second column: at 1100, nearly one player in five grabs on e5 with the knight, and one in seven takes the rook on a1 (which loses after 10.Qa4+ and the bishop is trapped). By 1900 the rook grab is down to 8% and the knight-to-e5 lunge has halved. The theory hasn’t changed. What changed is the population’s resistance to the obvious material grab.

That is the diagnostic. If you’re 1400 and you would have played Bxa1, you now know exactly two things: what the right move was, and that your instinct sits in the tail of the 1100 distribution rather than the 1900 one. Stockfish tells you the first thing only, and it tells you with an evaluation number that implies you should have seen a bishop trap eight plies deep.

Reading Maia’s output as a probability, not a verdict

Get comfortable with this framing, because it reorganises how you use every engine afterwards.

Stockfish output: +0.72. That is a claim about the position under perfect play from both sides. It is objectively true and practically irrelevant if neither you nor your opponent can play perfectly.

Maia output: a list like Nf5 34%, Rd1 22%, h3 15%, Qc2 9%. That is a claim about people. It says that if you showed this position to a hundred 1500-rated players, about 34 of them play Nf5.

Now put those two side by side and you get something neither gives you alone. Suppose Stockfish says the best move is g4, worth +1.4, and the second-best is Nf5 at +0.3. Meanwhile maia-1500 gives g4 a 3% policy share. You’ve just found a move that is objectively winning and that almost nobody at your level considers. That’s a blind spot, and it has a name in the literature. The Maia authors built exactly this comparison to identify what they called “human blunders”: positions where a large share of players at a given rating make a specific error that the engine sees instantly.

The inverse is the more useful training signal. Positions where Maia’s top move has 60%+ share and Stockfish rates it as a clear mistake are traps. You are statistically likely to fall into them, and so is everyone you play. Those positions are worth memorising in a way that a random Stockfish blunder-check is not.

Turning this into an actual training routine

Three concrete uses, in increasing order of effort.

Play the Maia bots on Lichess. The accounts are live: maia1, maia5 and maia9 (corresponding to the 1100, 1500 and 1900 networks). They accept challenges, they play instantly, and they’re free. Pick the network 200-300 points above your rating and play twenty games. You will lose most of them, but the losses will be losses to positional squeeze and tactical patterns you recognise, not to engine-flavoured alien moves. Losing to a bot that plays like a good version of you is diagnostically rich.

Run the “double-check” analysis on your own games. Load a game into a GUI that supports multiple engines (Nibbler and En Croissant both handle this cleanly; Nibbler was built for Lc0 networks and is the easier starting point). Analyse once with Stockfish, once with maia-X where X is your own rating. Then export both and look for the divergence: every position where Maia’s top move differs from Stockfish’s top move by more than about 0.5 pawns. In a 40-move game you will typically find between four and eight such positions. Those eight positions are your entire study plan for the week. Not the fifty moves where Stockfish shaved 0.1 off your eval.

Use rating-ladder comparison on your recurring mistakes. When you find a mistake you make repeatedly, feed the position to maia-1100 through maia-1900 in sequence and watch the policy share of your bad move. If it decays smoothly from 30% at 1100 to 4% at 1900, you’re looking at a skill that improves with general rating and you’ll grow out of it. If the share stays flat at 20% across the whole ladder, you’ve found something that strong club players get wrong too, which usually means it’s a genuine pattern worth drilling rather than a symptom of general weakness.

Where Maia stops being useful

Maia is not an analysis engine and will happily hand you a losing move with high confidence. Its playing strength is roughly its training band, meaning maia-1900 is about a 1900. Ask it to evaluate a queen endgame and it will produce something a 1900 would produce, which is to say wrong. Never use Maia to decide whether a move is good. Use it to decide whether a move is likely, and let Stockfish handle good.

There’s also a coverage problem. Maia’s networks were trained on Lichess blitz and rapid games from around 2017-2019. Positions that essentially never occur in that pool (deep theoretical novelties, seven-piece tablebase territory, weird composed studies) get garbage predictions, because the network has no human behaviour to imitate. The policy distribution goes flat and meaningless. If Maia’s top move sits at 8% share with nothing else above 6%, treat the output as noise.

And it is one tool among several with very different jobs. Stockfish 17’s NNUE evaluation gives you ground truth, Leela gives you a positional second opinion built from self-play rather than human imitation, and Maia gives you the behavioural layer. Our engines compared guide walks through how each of the four architectures actually differs under the hood and when to reach for which.

The reframe

The reason to care about all this isn’t that Maia is stronger or smarter. It is neither. It’s that for the first time you have an engine whose errors are data about you rather than artefacts of a handicap slider.

Every other tool in your kit answers “what is the truth about this position?” Maia answers “what does a person like me do here?” Those are different questions and the second one is the one standing between you and 200 rating points. You already know your moves are worse than Stockfish’s. What you don’t know is which of your moves are worse than the moves of players 300 points above you, and which are already fine.

Start there. Load maia-1500, play ten games, and count how many times you thought “yeah, I’d have played that too.”