• 4 min read
Gemini beats ChatGPT as a D&D Dungeon Master
A D&D player tests Gemini 3.6 Flash and ChatGPT 5.5 as Dungeon Masters. Gemini wins, but neither matches an in-person game.

Image: TechRadar
Dungeons & Dragons is built around friends, dice, and the unpredictable decisions people make at a table. That human element is exactly what makes the game difficult to reproduce—but general-purpose chatbots already know the rules of D&D 5e and are being used by players who cannot, or do not want to, assemble a group.
To see how close that experience has become to a real game night, I tested Gemini 3.6 Flash and ChatGPT 5.5 as Dungeon Masters. Both received the same character and the same instructions: run a classic high-fantasy dungeon crawl, use a strict random-number generator for rolls, and roll dice for the player when necessary.
The test: a 10th-level conjuration wizard
The prompt began, “Can you be my D&D 5e DM?” It was followed by a request for a classic high-fantasy dungeon crawl and a full character sheet for a 10th-level conjuration wizard.
That choice was deliberate. A first-level thief would present relatively few challenges, while a high-level wizard can teleport, summon creatures, and use powerful spells. Both chatbots initially created a similar setting: an abandoned dwarven stronghold filled with magical traps. Their handling of the adventure quickly diverged.
Gemini 3.6 Flash runs a recognizable dungeon crawl
Gemini delivered the more conventional D&D experience. The wizard avoided a pressure-plate trap that fired darts, then defeated skeletons by casting Grease on a stone staircase and freezing them with Ray of Frost as they slipped toward the bottom.

Recommended reading
Smart glasses face a privacy test
At first, Gemini offered suggested actions after each scene. Those options made the game easy to play, but they also encouraged passive clicking. After I asked Gemini to remove them, I had to describe each action myself, which required more thought and made the interaction feel closer to making decisions at a table.
The adventure ended with a secret passage and a fight against a Naga, a giant snake monster. The wizard teleported away before attacking with Cone of Cold, freezing the creature solid. Checking the Naga’s hit points against the official book showed that Gemini had used the correct figure.
Every roll in the adventure succeeded, raising the question of whether Gemini had adjusted the results. Gemini said it used a pseudorandom number generator, rather than a true random-number generator, but denied changing the results to make the game easier or longer. Its writing was natural and occasionally evocative, although still weaker than a skilled human Dungeon Master.
ChatGPT 5.5 turns the crawl into a mystery
ChatGPT used the same character and broadly similar setting, but produced little or no combat during the entire session. It also did not tell me when it was rolling behind the scenes, choosing instead to preserve the narrative flow.
When asked why the adventure contained so little rolling, ChatGPT acknowledged the problem:
“That is a fair critique. The short answer is: I leaned too heavily into narrative pacing and not enough into D&D mechanics. The adventure became closer to a collaborative fantasy novel than a 5e dungeon crawl.”
The result was a mystery about a dimensional prison beneath the mountain, centered on a sphere suspended in a cavern. Its guardians delivered ominous statements such as, “The archive does not open for authority. It opens for judgement.” The wizard then investigated the prison’s origins and collected obsidian keys and seals while receiving more cryptic warnings.
The structure had the surface features of a mystery, but not the carefully planted clues and payoff expected from one. After 40 minutes, the session had produced very little conventional action, and the increasingly portentous narration became distracting.
ChatGPT also repeatedly used a familiar construction: “This isn’t just X. It’s Y.” Examples included, “Unlike the rest of Khaz Varek, this room has remained untouched. No dust. No decay,” and, after the wizard met a guardian: “You hear it too. Not a question. A confirmation.” The repeated pattern made nearly every statement sound like a major revelation.
Gemini wins, but neither replaces a group
Between the two, Gemini 3.6 Flash was the clear winner. Its session was mildly enjoyable, mechanically closer to a D&D dungeon crawl, and worth revisiting. ChatGPT’s adventure was unenjoyable, and the test found it worse than the author’s previous experience with the chatbot.
Both systems understood the rules well enough to support a powerful character and create a playable setting. Neither reproduced the social experience of people gathered around a table. For players who cannot find a group, Gemini was the better option; for those who can, the author considers the Dungeon Master’s role safe from replacement for a long time yet.
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.
via TechRadar


