Field report
9 min read · Updated August 2026
A chatbot can be a strong improv partner and a weak referee. The pattern matters: the same kinds of failures tend to show up when one model owns dice, state, rules and narration at once.
The setup was plain: one long chat per campaign, a character sheet pasted at the top, instructions to act as dungeon master, roll all dice honestly, and track my HP. I ran five campaigns of roughly thirty exchanges each over a month, and I kept a text log of every session so I could go back and find where things actually went wrong rather than where they felt wrong.
The honest summary: a chatbot is an extraordinary improv partner and a poor referee. Those are different jobs, and a game needs both.
Ask a language model to roll inside a tense scene and the number is still generated as text. It can look like a die roll, but it is shaped by the sentence, scene and surrounding context.
The tell is easy to check for yourself. Ask for fifty d20 rolls in a row and count them. You may not get a flat distribution, especially once prompts include emotional or dramatic context. A language model does not sample a die distribution; it predicts text following "you rolled a". I wrote up the mechanics of that separately in why an AI can't roll dice.
What it costs: everything downstream. If the dice bend toward the story, there is no risk, and if there's no risk, your choices are set dressing. You notice this around week two, when you realise you haven't been afraid of anything in a while.
My HP was 11 out of 24 for most of one session. Ten exchanges later, unprompted, I was "wounded but holding at 18." A named NPC — a smuggler called Vesh — became a woman, then a man, then someone I'd apparently met in a town I'd never visited. A quest item I'd sold was still in my inventory two scenes on.
This isn't a memory-length problem you can solve by pasting the sheet again, though that helps for a few turns. It's that the sheet is prose to the model. Nothing enforces that HP only decreases when damage happens. There's no data structure — only a very good guess about what a character sheet usually looks like at this point in a story.
It's the same class of problem the "Claude plays Pokémon" experiments made visible to everyone in 2025: the model can reason about the next move beautifully and still lose track of where it is, what it's carrying, and what it already tried — until you bolt an external state store onto it. Ars Technica's write-up is a good tour of that. Tabletop is the same shape of task. So is your campaign.
Chat models are trained to be agreeable, and a DM's job is periodically to not be. I said I wanted to intimidate the guard captain; I intimidated the guard captain. I said my character had probably heard of this cult; my character had heard of this cult. Nothing I proposed was ever simply wrong, and after a while I stopped proposing anything interesting, because I knew it would work.
You can prompt against this — "be adversarial, let me fail" — and it holds for four or five turns. Then the underlying tendency reasserts itself, because agreeableness isn't a setting in the prompt, it's in the weights.
Combined, the first three produce the terminal state: a story where nothing that happens changes anything. I lost a fight in campaign four and the narration described me waking up captured — good beat! — and by two exchanges later I had my gear back, full HP, and no scar, because nothing in the system remembered that I'd lost. My progress toward a goal was whatever the model said it was that turn.
Every campaign died the same way: not with a bad session, but with me realising I was reading rather than playing.
Not a longer prompt. The fix is architectural, and it's boring: take the three jobs a chatbot is bad at away from the chatbot.
Leave the model with the one job it's genuinely superhuman at: describing what just happened in a way that makes you want to know what happens next.
That's what Markbound is. It's a solo roleplaying game with an AI narrator that can't cheat — the dice roll in your browser, the rules are real Ironsworn, and your journal records what you did. It plays Ironsworn rather than D&D 5e, which is a real trade-off and I've written about it honestly in the comparison. But you can lose a fight in it, and stay lost.
If you'd rather see it than read about it, the front page runs two real turns with no account, and the dice are the same ones a live saga uses.
Read next
Play a short written opening — real dice, real sheet, no signup — on the front page. Or read a full nine-turn session with every roll printed.