Yes, and better than most people expect — right up until it doesn't. The interesting part is which half of the job it is good at, and what a program has to take away from it before the other half works.
A game master does two jobs that have almost nothing in common. One is imaginative: describe the room, play the frightened clerk, decide what the noise upstairs was. The other is administrative: hold the numbers, apply the rules, remember what you were told in session three, and keep a secret for four hours without leaking it.
A language model is startlingly good at the first job and structurally bad at the second. Nearly everything worth knowing about AI game masters follows from that one split.
Every one of these will happen to you in a plain chat window within an hour. They are not model-specific and they are not fixed by better prompting alone.
Models are trained to be helpful, and in a story helpfulness reads as: you find the clue, the guard believes you, the door was unlocked after all. Ask whether the ritual works and it will lean towards yes. Push back on an outcome you did not like and it will very often simply revise it. A referee who cannot disappoint you is not running a game, and horror in particular is entirely made of things not going your way.
What we do: the model never decides an uncertain outcome. It says what is being attempted and the program rolls, against a number from your sheet, in front of you. It then narrates the result it was handed. It cannot soften a failure it did not choose, and it cannot be argued into a different one.
Ask a model to track hit points across two hours and they will drift. Not through malice — arithmetic inside a stream of prose is simply not what the machinery is for. The same applies to ammunition, to how many doses are left, and to what a skill was rated at before the injury.
What we do: the program owns every number on the sheet — health, sanity, luck, equipment, the companions’ sheets too. The Keeper is told what the numbers are and reports what should change; it never holds them. Wounds are applied by code, not by agreement.
A conversation has a finite window. Past a certain length the beginning falls out of it, and the failure is silent: the model does not say “I no longer remember the letter you found”, it simply carries on as though there had never been one. In a mystery, where the whole pleasure is that early details turn out to matter, this is fatal.
What we do: the campaign is stored outside the conversation. Scenario notes, the case board, the handouts you have been given, the plot beat you are on, your notes and the Keeper’s are all held by the program and put back in front of the model on every turn. Older narration is folded into a running summary rather than dropped. The model is not asked to remember; it is reminded.
A referee is frightening because they know the ending. A model improvising freely does not know the ending — it is inventing it as it answers, which means the mystery has no solution until you ask about one, and clues cannot really point anywhere because nothing has been decided for them to point at. It reads fine for twenty minutes and hollow by the third session.
What we do: a case is prepared before it is played. Whether you bring a scenario or ask for an original one, it is worked into structured notes first — what is true, who knows it, what is actually happening — and the Keeper runs from those notes. The ending exists before you go looking for it.
You can, and it is a genuinely good evening the first time. What you are getting is the imaginative half without the administrative one — so it is at its best in the first hour, when nothing has to be remembered, no numbers have accumulated and no clue has had to pay off yet. The drop-off after that is not a prompting problem. It is the four things above arriving at once.
The useful way to think about a program like this one is that it is not an AI game master. It is a game that uses a model for the part models are good at, and refuses to let it near the rest.
Being fair about the other side: a model does not know you. A human referee reads that you are bored and changes the weather, notices the thing you found funny and brings it back an hour later, and pitches the whole evening at the person in the room. None of that is on offer here, and a good table is still better than this. What is on offer is a game tonight, at your pace, that remembers everything.
Midnight Mythos is a solo Lovecraftian investigation run this way — the Keeper narrates and plays the world, the program owns the dice and the record. How to play describes a session in detail, and solo roleplaying covers the wider practice, including the three ways of doing it that need no AI at all.