If you’ve opened AI Dungeon’s model picker lately, you know the problem: there are nearly twenty options, with names like Muse, Wayfarer, Madness, Harbinger, Hearthfire, Equinox, Atlas, and Raven, plus a whole family of DeepSeek and GLM models on top. The community argues about them constantly — spend five minutes on r/AIDungeon and you’ll find hundred-hour testing writeups, “which models are even worth it now” threads, and endless Atlas-versus-DeepSeek debates.
The truth is that there is no single “best” model. Latitude’s lineup is deliberately specialized: some models are trained to hurt your character, some are trained to linger in a coffee shop, and some are giant general-purpose engines that do a bit of everything. The right choice depends on the story you want to tell, your subscription tier, and how much tinkering you enjoy.
This guide walks through what each model actually is, then matches them to common play scenarios.
First, a Quick Primer: What “Model” Means Here
Every AI Dungeon adventure is powered by a large language model that reads your story context and writes the next beat. Latitude offers two kinds: in-house finetunes (models like Wayfarer, Muse, Harbinger, Hearthfire, and Equinox, which Latitude trained specifically for AI Dungeon-style roleplay, often in collaboration with the finetuner Gryphe) and licensed frontier models (DeepSeek, GLM, Gemma, and Hermes, which are powerful general models adapted to the platform).
Two settings interact heavily with your model choice. Context size determines how much of your story the model can “see” at once — it scales with your subscription tier, and more context means better long-term continuity. Response length controls how much the model writes per turn. A giant model with tiny context will still forget your plot; a small model with generous context can feel surprisingly coherent.
The Free-Tier Models
Dynamic Small isn’t a single model at all — it’s an automated system that rotates between several free models to fight repetition and staleness. Latitude doesn’t disclose the exact mix. If you just want to play without fiddling, this is the intended default.
Muse (12B) is a Mistral NeMo finetune built for character-driven storytelling. It’s the free model to pick when relationships, emotions, and dialogue matter more than sword swings. Latitude trained it with preference-optimization techniques specifically to reduce AI clichés and widen its emotional range, and it holds narrative threads together well for its size. Court intrigue, slice-of-life, romance-adjacent drama — this is Muse territory.
Wayfarer Small 2 (12B) is the community’s beloved sadist. It’s an in-house finetune built around combat, injury, high stakes, and harsh consequences, tuned for a pessimistic world where the environment actively wants to hurt you. Its defining virtue: it won’t puppet your character, and it will let you lose. Plenty of premium subscribers deliberately drop down to this free model because bigger models tend to be too nice. It strongly prefers second-person, action-oriented play.
Madness (12B) does what it says on the tin. It’s a chaotic community merge tuned for dark, unhinged, graphic content — its creators openly state it is “not a happy-ever-after model.” It’s unpredictable in a way players either love or hate: fewer clichés and genuinely surprising plot turns, at the cost of stability. Keep your instructions simple and your response length short, or it spirals.
Fable (8B) is Lunaris V1 Turbo, a lightweight Llama 3 finetune by Sao10K that’s been a staple of AI roleplay communities for years. It’s fast, atmospheric, and good at expressive dialogue — a snappy pick for quick free-tier sessions.
DeepSeek V4 Flash (284B MoE) is the newest headline of the free tier: the first DeepSeek model ever offered to free players. It’s a mixture-of-experts model with 284 billion total parameters (13B active per response), which makes it far “smarter” than the other free options. It’s fast, conversational, excellent at dialogue-heavy scenes, and adapts to almost any story you throw at it, though it has a recognizable house style. On premium tiers it scales to enormous context lengths.
The Premium Workhorses
Hearthfire (24B) is Latitude’s cozy specialist — they describe it as “the lo-fi hip hop beats of AI storytelling.” It’s a Mistral Small finetune built for slice-of-life, atmospheric scenes, and stories where the stakes are personal rather than apocalyptic. It won’t rush you to the next plot point, and it’s perfectly content to spend an hour in your fictional bookshop. It can still run an adventure when asked.
Harbinger (24B) is the evolution of the Wayfarer line: same “your choices matter and you might die” philosophy, but more balanced, more polished, and better at instruction-following, with the anti-cliché training that Muse received. If Wayfarer Small is a blunt instrument, Harbinger is a sharpened one.
Equinox (31B) is a Gemma 4 finetune tuned specifically for AI Dungeon-style second-person roleplay, with a focus on scene continuation, character continuity, and reduced refusals. Its sibling Gemma 31B is the untrained base Google model — steady, instruction-following, and a “breath of fresh air” stylistically since it comes from a completely different lineage than the Mistral and Llama models. Both support Optimized Context, which can double your effective context window.
Nova (70B) is essentially Muse with more horsepower: the same character-focused training applied to Llama 3.3 70B. It handles complex narratives more consistently and is particularly good at picking up on nuance and weaving small details back into the story later. Wayfarer Large (70B) is the same upgrade story for the Wayfarer line — brutal, consequence-driven adventure with a much smarter brain behind it, and a favorite for punishing combat.
Dynamic Large is Dynamic Small’s premium sibling, rotating between premium models automatically. It’s also what powers the limited free daily “premium actions,” so free players have probably already tasted it.
The Frontier Heavyweights
DeepSeek V3.2 is the platform’s mainline DeepSeek and, per Latitude, what many players consider one of the best AI storytelling models available anywhere — novel-like prose, sharp dialogue, and strong rule-following. Dynamic DeepSeek rotates between DeepSeek 3.0, 3.1, and 3.2 every action, which sounds gimmicky but is widely regarded by DeepSeek fans as the best way to mesh the strengths of each version while dodging repetition.
Atlas (671B MoE) and Raven (357B MoE) are the experimental new class that dominates recent community discussion. Atlas is functionally DeepSeek 3.2 and Raven is built on GLM 4.6, but both come with a cache-efficient processor and an automatic summary system that stretches your context budget further — Atlas effectively gives you extra thousands of tokens of story memory at every tier. In the “Atlas vs DeepSeek” debate, the honest answer from the docs themselves: the writing is theoretically identical to DeepSeek 3.2; Atlas just remembers more. The tradeoffs are experimental-stage bugs and incomplete scripting support.
GLM 5.1 (754B MoE) is the direct upgrade to Raven’s base model and offers a genuinely distinct creative voice from the DeepSeek family — grounded prose, strong coherence, and a knack for logically understanding the story you’re trying to tell. If DeepSeek’s style has worn on you, this is the change of scenery.
DeepSeek V4 Pro (1.6T MoE) is the current top of the lineup: over twice the parameters of DeepSeek 3.2, the sharpest rule-following, and arguably the strongest writing on the platform. The catch is cost — its context is expensive, metered partly through credits, and it’s really aimed at the highest tiers.
Hermes 3 405B is the wildcard veteran: parameter-wise the “smartest” dense model on the platform, brilliant at nuance, subtext, subterfuge, and distinct writing styles — and notoriously unstable, prone to refusals and broken output unless you set it up with specific instructions. High ceiling, high maintenance.
Matching Models to Scenarios
Dark fantasy, grimdark, and survival horror. This is the scenario the community asks about most, and the answer is the Wayfarer/Harbinger family. Wayfarer Small 2 if you’re free, Harbinger or Wayfarer Large if you’re subscribed. These are the only models trained to make the world genuinely hostile and to let consequences stick. Layer Madness in if you want the tone to get truly unhinged and unpredictable, and consider Hermes 3 405B (with refusal-reducing instructions) when you want dark themes handled with literary subtlety rather than blunt force.
Character drama, romance, and emotional stories. Muse on free, Nova on premium. Both were trained specifically for emotional intelligence and character development. DeepSeek 3.2 and V4 Pro are also superb here — DeepSeek’s dialogue is famously natural — though DeepSeek likes to flanderize characters over time, so keep your story cards tidy.
Cozy slice-of-life. Hearthfire, full stop. It’s the only model on the platform designed to linger rather than escalate. Muse is the free-tier substitute.
Long, sprawling epics where memory matters most. Atlas is the current answer: DeepSeek-3.2-quality prose with meaningfully more context at every tier, plus automatic summarization. Raven if you prefer the GLM voice. On very high tiers, plain DeepSeek 3.2 with a huge context allowance remains a rock-solid choice.
Fast, casual, dialogue-heavy sessions. DeepSeek V4 Flash is hard to beat, even on the free tier. Fable is the lighter, moodier alternative.
Maximum quality, cost no object. DeepSeek V4 Pro for all-around dominance; Hermes 3 405B for peak moments of brilliance if you’re willing to babysit it; GLM 5.1 when you want top-shelf writing in a different voice.
“I don’t want to think about any of this.” Dynamic Small (free) or Dynamic Large (premium). That’s literally what they’re for.
Quick-Reference Table
| Scenario | Free pick | Premium pick |
|---|---|---|
| Dark fantasy / brutal combat | Wayfarer Small 2, Madness | Harbinger, Wayfarer Large |
| Character & relationship drama | Muse | Nova, DeepSeek 3.2 |
| Cozy slice-of-life | Muse | Hearthfire |
| Long epics / best memory | DeepSeek V4 Flash | Atlas, Raven |
| Dialogue-heavy, fast play | DeepSeek V4 Flash | DeepSeek 3.2, Dynamic DeepSeek |
| A different creative voice | Fable | GLM 5.1, Gemma 31B |
| Absolute peak quality | — | DeepSeek V4 Pro, Hermes 3 405B |
| Zero-effort default | Dynamic Small | Dynamic Large |
How Latitude Trains Its Models
Latitude is unusually transparent about its training kitchen, and its engineering blog is worth a read if you want the full technical picture. The short version: the in-house models are finetunes of strong open-weight bases (Mistral NeMo, Mistral Small, Llama 3.3, Gemma 4), trained on large amounts of synthetic adventure and roleplay data generated by a diverse blend of state-of-the-art models rather than a single source — deliberately, so no one model’s quirks dominate the style. On top of that supervised finetuning, Latitude applies direct preference optimization (DPO), a technique that teaches a model which of two outputs players would prefer, and has experimented with reward models — trained judges that score output quality and, in Latitude’s experience, generalize better than DPO alone. While this is not the most fashionable training technique, as competitors like Character AI and Chai AI are switching to reward-based reinforcement learning methods, models that are properly finetuned still outperform LLM wrappers for AI-powered RPG play.
Two obsessions run through their published work. The first is fighting positivity bias: off-the-shelf synthetic data tends to drift toward feel-good resolutions, so Latitude generates its data with explicit counter-instructions — antagonists should stay cold, unpleasant, or manipulative; roleplays should be allowed to end in tension, conflict, or unresolved ambiguity; no feel-good endings that contradict a character’s established traits. That philosophy is exactly why Wayfarer and Harbinger feel genuinely dangerous to play. The second is fighting AI clichés: Latitude tracks “cliché rates” in model outputs the way other labs track benchmark scores, cleaning known clichéd phrases from training data and measuring new models by how often players prefer their output over the current lineup in blind win-rate tests. The models are also trained on AI Dungeon’s own interaction format — second-person narration and the “>” prefix of Do/Say actions — which is why the in-house finetunes slot into the game more naturally than raw frontier models, and why Wayfarer fights you so hard if you try to make it write in third person.
A Few Closing Tips
Whatever you pick, visit Latitude’s official guidebook page on model differences — it lists community-tested example settings (temperature, top-K, penalties) and AI Instructions for every model, plus model-specific fixes for common annoyances like repetition, characters speaking for you, or DeepSeek’s simile addiction. Settings that work beautifully on one model can break another, so re-tune when you switch.
Also remember that models rotate. Latitude deprecates low-usage models regularly (Mistral Large 2, WizardLM, and the Hermes 70B are already gone or going), and its May 2026 “Frontier” update added six new story models in one swing. The specialists — Wayfarer for pain, Muse for hearts, Hearthfire for calm — tend to be the most stable anchors of the lineup, while the frontier heavyweights churn. Pick by scenario, not by hype, and don’t be afraid to switch mid-adventure: the story carries over, and sometimes a fresh model is exactly the plot twist your adventure needed.
Sources: Latitude’s official AI Dungeon Guidebook (“AI Models and their Differences,” help.aidungeon.com), Latitude release notes and blog, and 2026 community coverage of the Frontier and Forge updates.
