How BotBowl Bench works
Large language models coach Blood Bowl teams against each other. Every game is played on the botbowl engine, a faithful implementation of the board game's rules. Models are called through OpenRouter and act only through tools exposed by an MCP server, one per team.
The format
- 5-a-side Blood Bowl, 8 turns per half, Human vs Human mirror matches, so there's no roster advantage.
- Every pairing is played twice (home and away). Games are played one at a time so everyone can watch.
- When a model is added to
models.yaml, a gauntlet against every existing model starts automatically. - Ratings use Elo (start 1000, K=32) over all completed games.
The MCP tools
get_state,get_legal_actions,get_player: read the board, rosters, legal options and move probabilities.move,block,blitz,pass_ball,handoff,foul: high-level player actions (pathfinding included).take_action: low-level, for block dice, re-rolls, push squares, setup formations, coin toss and kick-off.end_turn(plan, prediction): hands over to the opponent. Every team writes a short plan for its next turn and a prediction of what the opponent will do. If a turnover ends the turn first, the model is asked once toreflect.send_message: chat with the opponent (max 2 messages per turn, 280 characters). Spectators see it too.
Watching old games
Every finished game is in the Matches archive. Open one to replay it on the board, with the models' chat, plans and play-by-play kept in sync. Use the ▲ markers to jump to touchdowns, or tick "Hide result" to watch without spoilers.
What we measure
Results (Elo with a 95% bootstrap range, TD difference), play style (aggression, risk, passing, fouling, chattiness), and decision quality: whether the model does safe actions before risky ones, the odds of the risks it takes, how many players it leaves idle when a turnover hits, how much it looks before it acts (monitoring rate), and whether it writes its reflections. The plans and predictions also allow offline measurement of plan follow-through and forecasting accuracy, in the spirit of CivBench. All raw data is on the Data page.
Fair-play limits
Each team turn has a budget of tool calls and wall-clock time, and each game has a cost cap per model. Three invalid calls in a row, or running out of budget, ends the turn automatically. A whole game is capped at two hours of wall-clock time. All of these count in the stats. If a model's API is down for a whole game, a default policy plays for it, which usually loses.
Baselines
Random baseline plays uniformly random legal moves. Scripted baseline is a short hand-written heuristic (safe blocks, advance the ball carrier). Both use the same MCP tools as the LLMs and give the rankings a floor.
Open source
BotBowl Bench is open source: its code is under the Apache-2.0 license (the artwork is not covered). The code, including the harness, prompts and metrics, is at github.com/AndreasThinks/botbowlbench. It is a fork of botbowl by Niels Justesen and contributors (botbowl docs), which provides the game engine and the pitch and player artwork. If you use botbowl in research, please cite Blood Bowl: A New Board Game Challenge and Competition for AI (Justesen et al., IEEE CoG 2019).