Sparring Against Human-Like Bots At Your Rating
Most improvers between 1000 and 1900 have the same setup: Stockfish is one click away on the Lichess analysis board, the eval bar swings around during review, and none of it changes what happens in the next game. The engine is not the problem. The problem is that a 3600-rated oracle is a terrible sparring partner and an even worse teacher unless you know exactly what to ask it. Human-like chess bots to play against fix the first half of that. Knowing how to interrogate an engine afterwards fixes the second half.
This page covers both: which bots actually resemble a player at your strength, how to run them locally so you get unlimited games from any position, and how to turn Stockfish’s raw UCI output into a list of things to work on that doesn’t just say “you blundered a knight on move 23.”
The 1500 That Isn’t A 1500
Open Stockfish, set Skill Level to 5, and play it. You will lose to a series of moves that no 1500 has ever played, and you will win a game where it hangs a rook on move 9 for no reason. That’s the mechanism working as designed: Skill Level weakens Stockfish by making it pick a random move from its top candidates with a probability tied to the level, and by capping search depth. The result is a strong player with a seizure disorder. Its positional understanding is still superhuman; its errors are uniformly distributed noise.
UCI_LimitStrength plus UCI_Elo is a real improvement over Skill Level, and in Stockfish 17 the range bottoms out at 1320. It’s calibrated against engine-vs-engine results, though, not against human games. A “1500” Stockfish will still find a six-move tactic more reliably than a 2000 and then drop a piece to a one-move fork, because the thing that gets throttled is search, and search is not what separates a 1200 from a 1700.
Here is what that costs you in training terms. When a real 1500 goes wrong, they go wrong in patterned ways: they take on a bad bishop trade to “develop,” they push a pawn to gain space and create a hole, they miss a backwards defensive move, they start the wrong plan in a structure they’ve seen twice. Those patterns are what you need to learn to punish, because those are the errors you’ll face on Saturday morning. Random noise teaches you nothing except to check for hanging pieces, which you should be doing anyway.
What Makes A Bot Human-Like
Maia is the project that actually solved this. It came out of a collaboration between researchers at Toronto, Cornell and Microsoft Research, published at KDD 2020 under the title Aligning Superhuman AI with Human Behavior: Chess as a Model System. Rather than training a network to win, they trained it to predict the move a human of a specific rating would play, using roughly 12 million Lichess games per rating band.
There are nine models: maia-1100.pb.gz through maia-1900.pb.gz, in 100-point steps, each trained only on games where both players sat in that band. They run as Leela Chess Zero weights, but with a crucial difference in how you invoke them. A normal Leela run does thousands of MCTS rollouts. Maia is designed to run at one node, meaning you get the raw policy head output with no search on top. No search is the point. The network has learned what a 1500 plays, and search would only make it better than a 1500.
The move-matching numbers are the part worth internalising. Maia matches the actual human move just over 50% of the time, peaking around 52% when the model band matches the player band. Stockfish-derived and standard Leela baselines in the same paper top out closer to 42%, and their accuracy is flat across ratings: they play the same way against a 1100 and a 1900. Maia’s accuracy curve has a peak, which is the signature of having learned a specific skill level rather than a general one. It also predicts which moves a human will blunder, at rates well above chance.
What this means at the board: Maia-1500 will develop naturally, castle at a sensible moment, put a rook on a half-open file, and then walk into the same pin structure you walk into. It will find the obvious two-move tactic and miss the quiet one. It resigns badly and defends worse than its rating in dead-lost positions, because losing positions are underrepresented in its training data (humans resign). Those are real limitations, and I’ll come back to them.
Choosing The Band You Actually Need
On Lichess the three deployed accounts are maia1 (1100), maia5 (1500) and maia9 (1900). Challenge them like any other player. Don’t be thrown by the rating displayed on the account profile, which drifts hundreds of points above the nominal band because these bots face a heavily skewed pool of challengers. The training band is the honest signal, not the Glicko number on the page.
For a wider survey of what else is available, including the Leela odds bots and a few oddities worth a game each, see The Best Lichess Bots To Play Against At 1000-1900. The short version of the selection logic:
| Your rating | Main sparring band | Stretch opponent |
|---|---|---|
| 1000-1250 | maia-1100 | maia-1300 |
| 1250-1500 | maia-1300 | maia-1500 |
| 1500-1700 | maia-1500 | maia-1700 |
| 1700-1900 | maia-1700 | maia-1900 |
Play roughly 70% of your sparring games at your own band and 30% one or two bands up. Same-band games tell you whether your decisions beat the decisions of your actual peers. Games against a higher band are for punishment: they show you which of your habits survive contact with someone slightly better.
LeelaQueenOdds, LeelaRookOdds and LeelaKnightOdds on Lichess are a different tool entirely. They’re Leela networks trained specifically to win while down material against humans, and they are brutally good at it. Playing a piece up against something that expects to be a piece down is the best defensive-technique drill I know of, because the usual “I’m winning, I’ll just trade everything” plan gets punished immediately.
Chess.com’s side of this is the personality bot roster (Martin at 250 is the famous one) plus the adaptive bot that tracks your rating. They’re engine-based with move-selection rules layered on, so the errors are less patterned than Maia’s, but the bots do think for a variable amount of time, which Maia does not. If you want an opponent whose clock behaviour feels vaguely human, that’s where to get it.
Running Maia On Your Own Machine
Playing on Lichess gets you games. Running Maia locally gets you positions, which is what you actually need for drilling. Grab an lc0 binary from lczero.org and the weights from the CSSLab maia-chess repository on GitHub.
The flags that matter:
lc0 --weights=maia-1500.pb.gz \
--backend=eigen \
--minibatch-size=1 \
--max-prefetch=0 \
--threads=1
Then, in whatever GUI you’re using, set the engine to a fixed node count of 1. Not depth, not time. If you give it time, lc0 will search and you’ll be playing something considerably stronger than 1500. In Nibbler this is straightforward: Nibbler is Leela-native and shows you the policy percentages for each candidate move, which is genuinely interesting on its own (you get to see how confident the 1500-model is that a 1500 plays Nf3 here). En Croissant will happily manage both Maia and Stockfish side by side, and it handles playing from a FEN well.
One quirk to know: at one node, Maia is deterministic. Same position, same move, every time. That kills variety in the opening but it’s a gift for drilling, which I’ll get to. If you want variety, lc0’s --temperature flag samples from the policy distribution instead of taking the argmax. Setting --temperature=0.4 gets you noticeably different openings while keeping the move quality recognisable; push it to 1.0 and it starts playing things a 1500 would only play by accident.
If you want to verify what you’ve got, run a match with cutechess-cli:
cutechess-cli \
-engine name=maia1500 cmd=lc0 arg=--weights=maia-1500.pb.gz \
arg=--minibatch-size=1 arg=--max-prefetch=0 nodes=1 \
-engine name=sf1500 cmd=stockfish \
option.UCI_LimitStrength=true option.UCI_Elo=1500 \
-each proto=uci tc=60+0.6 -rounds 100 -pgnout calib.pgn
Then open calib.pgn and look at how each side lost. The score will be somewhere near even. The games will not look remotely similar, and that difference is the whole argument for using Maia.
Interrogating Stockfish Instead Of Obeying It
Now the analysis half. Almost every club player uses Stockfish in single-PV mode, watches the bar, and reads off the top move. That’s the least informative way to use it.
Three settings change everything. In a UCI-capable GUI or straight in a terminal:
setoption name MultiPV value 4
setoption name UCI_ShowWDL value true
setoption name Threads value 8
setoption name Hash value 2048
position fen r2q1rk1/bpp3p1/p1nppn1p/4p3/4P3/2PP3P/PP1N1PP1/R1BQR1K1 w - - 0 12
go nodes 20000000
MultiPV 4 gives you the top four moves with separate evaluations, which converts “the engine move” into “here is the spread of reasonable plans.” UCI_ShowWDL is the underused one. go nodes instead of go depth or go movetime makes your analysis reproducible across machines and across sessions, which matters if you’re keeping notes.
Output comes back looking like this:
info depth 34 seldepth 45 multipv 1 score cp 31 wdl 168 806 26 nodes 20001233 ... pv d2c4 ...
info depth 34 seldepth 44 multipv 2 score cp 24 wdl 141 832 27 nodes 20001233 ... pv d1b3 ...
info depth 34 seldepth 46 multipv 3 score cp 19 wdl 128 843 29 nodes 20001233 ... pv d3d4 ...
info depth 34 seldepth 43 multipv 4 score cp 14 wdl 112 855 33 nodes 20001233 ... pv d2f1 ...
Exact centipawn values will differ by a few points on your build and hardware. The structure is what you’re reading. Four candidate moves inside 17 centipawns of each other means the position has no “engine move” in any meaningful sense, and anyone telling you they lost the game by playing the third-listed option is wrong.
The wdl field is per-mille win/draw/loss from the side to move, at engine strength. At cp 31 Stockfish is telling you it wins 16.8% of the time, draws 80.6%, and loses 2.6%. That’s a position it considers essentially drawn. Meanwhile the eval bar shows a visible white advantage, and a 1500 playing it as White scores well above 50%. Both things are true. The bar is calibrated to a player who doesn’t exist.
Centipawns Are Not Points
Lichess converts centipawns to a win percentage before deciding whether something was an inaccuracy, a mistake or a blunder. The formula is:
Win% = 50 + 50 * (2 / (1 + exp(-0.00368208 * cp)) - 1)
Worth having the shape of that curve in your head:
| Eval | Win% |
|---|---|
| +0.50 | 54.6 |
| +1.00 | 59.1 |
| +2.00 | 67.6 |
| +3.00 | 75.1 |
| +4.00 | 81.4 |
| +5.00 | 86.3 |
Lichess flags a drop of 10 percentage points as an inaccuracy, 20 as a mistake, 30 as a blunder. Run a few conversions through that and the practical consequences fall out immediately.
Dropping a clean pawn from dead equal, +1.00 to -1.00, is a swing from 59.1% to 40.9%: 18 points, a mistake, not a blunder. Going from +5.00 to +3.00, which feels catastrophic when you see it in review, is 86.3% to 75.1%: 11 points, an inaccuracy. But 0.00 to -2.00 is 50% to 32.4%, a 17.6-point mistake, despite being the same two-pawn swing. The eval bar’s linear-looking scale hides the fact that centipawns near zero are worth far more than centipawns near five.
That asymmetry should change where you spend review time. A player who bleeds three 12-point inaccuracies in an equal middlegame is losing more real winning chances than one who converts a +6 into a +3.5 and still wins. Review sessions that chase the biggest bar movements systematically ignore the first player’s problem, which is usually the bigger one.
A Worked Post-Mortem
Take the position from that UCI block, which arises after 1.e4 e5 2.Nf3 Nc6 3.Bc4 Bc5 4.d3 Nf6 5.c3 d6 6.O-O O-O 7.Re1 a6 8.Bb3 Ba7 9.h3 h6 10.Nbd2 Be6 11.Bxe6 fxe6. It’s a completely ordinary Italian, the kind you reach three times a week.
A 1600 sitting here typically plays 12.Nc4, because the knight has a square and knights go on squares. Then they discover the knight does nothing on c4, spend four moves rerouting it, and end up worse. When they run the game through Lichess afterwards, move 12 gets no annotation at all, because it cost 7 centipawns. The engine has told them, accurately, that nothing went wrong. Nothing evaluable went wrong.
The information you want is in the pv lines, not the scores. 12.Qb3 attacks e6 and b7 the instant the light-squared bishop leaves; 12.d4 challenges the centre while Black’s e-pawns are doubled and the dark-squared bishop on a7 is biting on granite; 12.Nf1 heads for g3 and then f5 or h5, with e6 as a long-term target. Three coherent plans, each scored within 17 centipawns of the others. 12.Nc4 is a fourth square with no plan attached, and a plan is the thing you were missing.
So the finding isn’t “12.Nc4 was inaccurate.” The finding is: in Italian structures where Black has recaptured on e6 with the f-pawn, I don’t have a plan, and there are three standard ones. That’s a note you can act on. Write it in a Lichess Study chapter with the FEN and the three pvs, and drill it against Maia from that exact position until you’ve played all three.
The Weekly Sparring Loop
Here’s a schedule that works, sized for someone with a few hours a week.
Play six games against your band and two against the band above, at 15+10 or 10+5. Long enough that your mistakes are decisions rather than mouse slips. Rapid is where your thinking actually lives.
Before you touch the engine, spend five minutes on your own post-mortem. Write down, in words, the move where you think the game turned, and what you were trying to do at that point. This step feels pointless and is the single highest-value thing in the loop, because it forces you to produce a prediction the engine can then falsify.
Now run the analysis with MultiPV 4. For every position where your win% dropped by 10 or more, record three things: the FEN, the category, and the spread of the top four moves. Categories that earn their keep at club level: missed tactic (theirs), missed tactic (mine), wrong plan, bad trade, time trouble, king safety, endgame technique. Keep the list short enough that you’ll actually use it.
After twenty games, count the categories. The distribution will be lopsided and it will probably surprise you. Lichess Insights gives you the same picture from a different angle: filter by phase and by opening and look at where your average centipawn loss spikes. A player who’s convinced their problem is openings usually finds their move-25-to-40 ACPL is double their move-10-to-20 number.
The final step is the one people skip. Take the five worst positions from those twenty games, load each FEN into your local Maia at your band, and play it out. Best of three from each side. You are not checking whether you now know the right move. You’re checking whether you can convert the resulting position against an opponent who plays like the people you actually lose to.
Drills That Use The Bot’s Determinism
Maia at one node repeats itself, and that turns it into a rehearsal partner rather than just an opponent.
Opening rehearsal is the obvious use. Pick your repertoire line, play it against maia at your band, and when it deviates from your preparation you’ve found the move your peers actually play (as opposed to the move the book says is critical). Play the resulting position five times. Because Maia is deterministic, you can iterate on your own choices while holding its responses fixed, which is the closest thing to a controlled experiment you’ll get in chess.
Endgame conversion drills work well too. Set up a rook and pawn versus rook, or a good-knight-versus-bad-bishop position from one of your own games, and convert it against the 1500-model. It will defend the way a 1500 defends: passively, with the rook in front of the pawn instead of behind it, missing the one active check that saves the game. Learning to beat that defence is more useful than learning to beat Stockfish’s defence, which you won’t manage anyway.
Then there’s the “predict the opponent” drill, which uses Nibbler’s policy display. Load a position from one of your games, set Maia to your opponent’s band, and look at the policy percentages before you look at Stockfish. If Maia-1500 gives 38% to the move that actually got played and 4% to the refutation you were worried about, you’ve learned something concrete about practical risk. Play the 4% line.
Where Bots Stop Helping
Maia has no clock. It moves instantly regardless of position, which means every game you play against it is a game where you learn nothing about time management, and time management is a substantial chunk of what separates a 1600 from a 1900. If your losses cluster in the last three minutes, sparring bots are not the fix, and you should be playing humans at the same time control instead.
Dead-lost positions are Maia’s other blind spot. Humans resign, so losing positions are scarce in the training data, and the models play them noticeably worse than their band would. Don’t use a bot to practise winning won games; the resistance isn’t real. The Leela odds bots are the correct tool for that, because they were trained specifically on the problem of fighting from a material deficit against a human.
And the engine itself has a limit that no amount of MultiPV fixes: it evaluates positions, not decisions. It has no opinion on whether you should enter a sharp line you’ll need to calculate accurately for eight moves, or a slow one where you’re comfortable. At 1400 the sharp line might be a blunder even when it’s objectively better, because you are the one who has to play it. Stockfish will never tell you that. The bot game, played at your own strength against something that goes wrong the way you go wrong, is what tells you.
In this section
The supporting pages under this subject.