Field report
9 min read · Updated August 2026
Not a hit piece. The first week was genuinely the most fun I'd had roleplaying alone in years. The month is what matters: the same four failures showed up in every campaign I started, at roughly the same point each time, and none of them are fixable with a better prompt.
The setup was plain: one long chat per campaign, a character sheet pasted at the top, instructions to act as dungeon master, roll all dice honestly, and track my HP. I ran five campaigns of roughly thirty exchanges each over a month, and I kept a text log of every session so I could go back and find where things actually went wrong rather than where they felt wrong.
The honest summary: a chatbot is an extraordinary improv partner and a poor referee. Those are different jobs, and a game needs both.
I asked for a d20 on a locked door. It gave me a 14 and the door opened. Later, in a fight I was clearly losing, I got three rolls in a row of 17, 19 and 18. That's not a run of luck — that's a model producing the number the scene wanted.
The tell is easy to check for yourself. Ask for fifty d20 rolls in a row and count them. You will not get a flat distribution. You'll get a hump around 12–17, almost no 1s, and a suspicious number of 20s at dramatic moments. A language model does not sample a distribution; it predicts the token most likely to follow "you rolled a". I wrote up the mechanics of that separately in why an AI can't roll dice.
What it costs: everything downstream. If the dice bend toward the story, there is no risk, and if there's no risk, your choices are set dressing. You notice this around week two, when you realise you haven't been afraid of anything in a while.
My HP was 11 out of 24 for most of one session. Ten exchanges later, unprompted, I was "wounded but holding at 18." A named NPC — a smuggler called Vesh — became a woman, then a man, then someone I'd apparently met in a town I'd never visited. A quest item I'd sold was still in my inventory two scenes on.
This isn't a memory-length problem you can solve by pasting the sheet again, though that helps for a few turns. It's that the sheet is prose to the model. Nothing enforces that HP only decreases when damage happens. There's no data structure — only a very good guess about what a character sheet usually looks like at this point in a story.
It's the same class of problem the "Claude plays Pokémon" experiments made visible to everyone in 2025: the model can reason about the next move beautifully and still lose track of where it is, what it's carrying, and what it already tried — until you bolt an external state store onto it. Ars Technica's write-up is a good tour of that. Tabletop is the same shape of task. So is your campaign.
Chat models are trained to be agreeable, and a DM's job is periodically to not be. I said I wanted to intimidate the guard captain; I intimidated the guard captain. I said my character had probably heard of this cult; my character had heard of this cult. Nothing I proposed was ever simply wrong, and after a while I stopped proposing anything interesting, because I knew it would work.
You can prompt against this — "be adversarial, let me fail" — and it holds for four or five turns. Then the underlying tendency reasserts itself, because agreeableness isn't a setting in the prompt, it's in the weights.
Combined, the first three produce the terminal state: a story where nothing that happens changes anything. I lost a fight in campaign four and the narration described me waking up captured — good beat! — and by two exchanges later I had my gear back, full HP, and no scar, because nothing in the system remembered that I'd lost. My progress toward a goal was whatever the model said it was that turn.
Every campaign died the same way: not with a bad session, but with me realising I was reading rather than playing.
Not a longer prompt. The fix is architectural, and it's boring: take the three jobs a chatbot is bad at away from the chatbot.
Leave the model with the one job it's genuinely superhuman at: describing what just happened in a way that makes you want to know what happens next.
That's what Markbound is. It's a solo roleplaying game with an AI narrator that can't cheat — the dice roll in your browser, the rules are real Ironsworn, and the story remembers what you did. It plays Ironsworn rather than D&D 5e, which is a real trade-off and I've written about it honestly in the comparison. But you can lose a fight in it, and stay lost.
If you'd rather see it than read about it, the front page runs two real turns with no account, and the dice are the same ones a live saga uses.
Read next
Two turns of real play — real dice, real sheet, no signup — run on the front page. Or read a full nine-turn session with every roll printed.