What’s Changed: Grok 4.5 is a coding and agentic model that became the default for everyone, including roleplayers. It scores 1521 Elo on coding and only 1440 on creative writing, and its hallucination rate more than doubled compared to Grok 4.3. Characters go flat, point of view drifts, and long stories lose continuity. You cannot roll back to an older model, but Project instructions and an anti-coding system prompt recover most of the quality.
If your Grok 4.5 roleplay got worse overnight, you are not imagining it, and you are not doing anything wrong. Characters that used to have teeth now answer like a helpdesk. Scenes that ran for pages wrap up in four sentences.
The confusing part is the timing. Grok 4.5 launched on July 8, 2026, but the complaints only exploded around two weeks later, which made a lot of people assume xAI had quietly nerfed something in secret.
There is a simpler explanation, and once you see it the rest of the problem makes sense too.
This piece covers what Grok 4.5 was built for, why that specific design choice hurts storytelling, whether you can go back to 4.3, and the exact settings that claw back most of the lost quality.

Why Is Grok 4.5 Roleplay Worse Than It Used to Be
Grok 4.5 roleplay is worse because the model was trained as a coding and agentic assistant, then made the default for every user.
Its Chatbot Arena Creative Writing Elo sits at 1440 against a Coding Elo of 1521, a gap that shows up directly in prose quality.

What is Elo: A skill rating built from head-to-head blind voting, where people pick the better of two answers. Higher means the model won more often.
xAI built Grok 4.5 on a 1.5-trillion-parameter foundation and co-trained it with Cursor on real developer session data. Cursor has said the training mix stayed deliberately broad, pulling in STEM tasks, research papers and other knowledge work rather than raw code alone.
That is still a model tuned to finish tasks precisely and stop. Roleplay wants the opposite behaviour, which is to keep going, take a risk, and stay in a voice that is not the assistant’s own.
The rollout timing explains the delayed reaction, and I think it is why so many people assumed a secret nerf. xAI held Grok 4.5 back from the European Union at launch, and EU users only saw it appear in their apps around July 23 and 24, which is when the complaint wave went vertical.
The version that people are comparing it against
Grok 4.1 is still widely treated as the high-water mark for creative work, praised for length, emotional range and worldbuilding. Grok 4.2 and 4.20 earned a similar reputation for writing long chapters without constant nudging.
By the time of the 4.3 release, users were already arguing that creative quality was sliding. Grok 4.5 landed on top of that existing frustration, which is why the reaction was sharper than the version bump alone would suggest.
| Version | Tuned toward | Creative reputation | Still selectable |
|---|---|---|---|
| Grok 4.1 | Emotional range, conversational warmth | Widely treated as the high point for fiction | No, retired May 15, 2026 |
| Grok 4.2 and 4.20 | Long-form generation | Wrote long chapters without nudging | No, retired May 15, 2026 |
| Grok 4.3 | General reasoning, 1M token context | Divisive, seen as the start of the slide | API only, as the redirect target |
| Grok 4.5 | Coding and agentic tasks, 500K context | 1440 creative Elo against 1521 coding | Default everywhere, no opt out |
Look down that last column if you write fiction with Grok. Each release narrowed the escape route, and this one closed it.
What Was Grok 4.5 Built to Do Instead
Grok 4.5 was built for agentic engineering work, and it is genuinely good at it.
It leads on coding and tool-calling benchmarks, prices at $2 per million input tokens and $6 per million output tokens, and runs faster than most models in its class.

Mainstream coverage of the launch focused almost entirely on that developer story. TechCrunch covered the launch as an Opus-class model aimed squarely at professional engineering work, which it is.
What none of that coverage examined is what happens when a model optimised for correctness becomes the only option for people writing fiction. I would argue that is the single most under-reported part of this launch.
There is a wrinkle worth knowing about the benchmark numbers too. Cursor disclosed that an earlier snapshot of its own codebase was accidentally included in the training data, which inflated Grok 4.5’s score on the CursorBench evaluation.
Why Did My Characters Lose Their Personality
Characters flatten because safety and instruction tuning overwrite the same parameters that carry stylistic nuance.
Research on the alignment tax describes post-training as something that tends to overwrite rather than build on general capabilities, producing partial forgetting of exactly the traits roleplay depends on.
Published research on mitigating alignment tax covers the mechanism in detail. A related finding describes final-layer decoding overriding fragile logic chains with generic, safe priors, which is close to a technical description of a character losing their edge halfway through a scene.
Two hard numbers make the trade-off concrete. Grok 4.5’s hallucination rate on the AA-Omniscience benchmark sits near 53.5%, up from roughly 25% on Grok 4.3, according to Artificial Analysis.
The context window also shrank. Grok 4.3 offered a one-million-token window and Grok 4.5 offers 500,000, which is the clearest explanation for long stories losing track of established details.
Here is what I would watch for, mapped to what is causing it.
| Symptom | Likely cause | Fix |
|---|---|---|
| Characters sound generic and polite | Assistant persona reasserting itself over your character | Move character rules into Project instructions |
| Replies switch to third person when you asked for first | Instruction-following favouring narration defaults | Add an explicit POV lock line to every scene |
| Story forgets details from earlier chapters | Context window cut to 500K tokens | Keep a running character and world file, re-paste it |
| Prose turns short and clipped | Task-completion tuning wrapping up early | Set a minimum length instruction, request book-style prose |
| Grammar breaks in non-English languages | Weaker multilingual generation in this release | Write scene instructions in English, request output language explicitly |
| Model treats your instructions as spoken dialogue | Metadata confusion in narrative context | Wrap instructions in brackets or a separate labelled block |
Is Grok 4.5 More Filtered or Just Weaker at Writing
Most of the “Grok 4.5 is censored now” complaints do not hold up under scrutiny, and the writing-quality complaints do.
Plenty of users report that restricted content still generates without much resistance, while nearly everyone agrees the prose itself got worse.
What I would separate here is permission from execution. Grok 4.5 will still write the scene you asked for, and it writes it without the craft that made earlier versions worth paying for.
That distinction matters for how you fix it. Fighting a filter that is not the real obstacle wastes your time and your weekly message budget.
There is a genuine counter-argument worth stating fairly. Some users report Grok 4.5 handles pacing and restraint better than 4.3 did, holds prompt details that other models drop, and improved noticeably over its first days of release.
That last point is not just perception. Several people testing the same scene on consecutive days found grammar errors disappearing and output quality climbing, which suggests xAI has been tuning the model in place since launch.
Can I Switch Back to an Older Grok Model
No, and this is the part that frustrates people most. The consumer app gives you auto, quick and expert modes without naming the underlying model, so there is no version selector to fall back to.
Earlier releases came with a beta window where you could compare the new model against the old one and give feedback before the swap became permanent. Grok 4.5 arrived as a straight replacement, which broke a pattern users had come to expect.
The API route is closed too. xAI retired eight legacy Grok models on May 15, 2026 at 12:00 PM PT, including the fast and reasoning variants that creative writers favoured.
Requests to those retired model names do not error out. They silently redirect to Grok 4.3 with reasoning effort locked, and they bill at Grok 4.3 rates of $1.25 per million input tokens and $2.50 per million output tokens.
What I find telling is that the forced-default behaviour was not limited to the consumer app. Cursor users filed a bug report on July 22, 2026 about the IDE force-enabling Grok 4.5 and overriding their selected model even when Grok 4.5 had been switched off in the settings.
How Do I Fix Grok 4.5 Roleplay Quality
The single highest-impact change is moving your character rules out of custom instructions and into Project instructions, which Grok weights more heavily.
From there, an anti-coding system prompt recovers most of the remaining quality.
Cache and settings tweaks are not the answer here, so here is the sequence I would work through in order:
- Create a Project for your story and put character sheets, tone rules and world details in the Project instructions rather than the general custom instructions box.
- Add an explicit anti-coding line to your instructions so the model stops applying task-completion behaviour to fiction.
- Lock the point of view in every scene opener, since this release drifts to third-person narration when left unspecified.
- Ask for book-style prose explicitly, which reliably pushes the model out of summary mode and into scene writing.
- Keep a separate character and world file, and re-paste it whenever a long story starts losing details, because the smaller context window fills faster than it used to.
- Ban reframing language directly, since the model likes to summarise and sanitise a scene rather than play it out.
The prompt wording matters more than usual on this release. Vague creative direction gets you the assistant voice back within two or three exchanges.
Before: “Write this scene in a more creative and interesting way, stay in character.”
After: “Optimise for advanced story-telling. Do not code unless specifically instructed to do so. Write in first person as the character, present tense, minimum 400 words. Describe what is happening like in a book. No reflective reframing or transformational framing. Analyse the character traits in the attached file and model all dialogue on those traits.”
That second version works because every clause blocks a specific failure this model has. The anti-coding line pushes back on the task-completion tuning, the POV lock stops the narration drift, and the anti-reframing clause stops it summarising your scene back at you.
One more thing to budget for. Expert mode and a SuperGrok subscription do produce longer and somewhat better replies, but they do not solve the underlying regression, and roleplay burns through Grok’s weekly usage limits fast enough that heavy users hit the cap in a few days.
What Should I Use if Grok Is Not Working for Roleplay
If you mainly used Grok as a writing partner rather than a coding tool, this release is a reasonable moment to stop fighting it.
The fixes above genuinely help, though they are effort you did not have to spend three versions ago.
For people running their own setup, open-weight models have closed much of the gap, and roleplay on DeepSeek is the route I would look at first for prose quality without a subscription.
If you want a companion that holds character without prompt engineering every session, a purpose-built platform is a better fit than a general model. Candy AI is built around persistent characters and long-term memory, so persona drift and forgotten details are not problems you have to manage yourself.
Nectar AI is the other one I would put in front of anyone coming from Grok, particularly if image generation inside the scene matters to you.
None of this makes Grok 4.5 a bad model. It is a strong engineering tool that happens to have been handed to an audience that wanted a novelist, which is worth remembering before you cancel anything.
The same complaint pattern showed up when Grok chat started feeling nerfed earlier this year, and some of that did get walked back.
Frequently Asked Questions
Is Grok 4.5 worse than Grok 4.3 for creative writing?
For most creative writers, yes. Grok 4.5 scores 1440 Elo on creative writing against 1521 on coding, and its hallucination rate roughly doubled to 53.5% from Grok 4.3’s 25%. Coding and agentic performance improved substantially.
Can I go back to Grok 4.1 or 4.2?
No. The consumer app offers only auto, quick and expert modes with no version selector. xAI retired eight legacy models from the API on May 15, 2026, and requests to those names now redirect silently to Grok 4.3.
Does SuperGrok or Expert mode fix the roleplay problem?
Partly. Expert mode produces longer and somewhat more considered replies, but the underlying regression remains. Roleplay also exhausts weekly message limits quickly, which makes continuous storytelling difficult on any tier.
Why does Grok 4.5 keep switching to third person?
This release defaults to narration when point of view is not locked explicitly in the prompt. Adding a POV instruction to every scene opener, rather than once at the start of a chat, fixes it in most cases.
Is Grok 4.5 more heavily filtered than earlier versions?
The evidence is mixed and mostly points the other way. Many users report restricted content still generates normally, and some find moderation more permissive than before. The consistent complaint is prose quality, not permission.
Why did my non-English roleplay break?
Grok 4.5 generates noticeably weaker grammar in non-English languages, with morphologically complex languages like Hungarian hit hardest. Writing your scene instructions in English while explicitly requesting output in your target language works better.
Quick Takeaways
- Grok 4.5 is a coding model that became everyone’s default, scoring 1521 Elo on coding against 1440 on creative writing.
- Its hallucination rate more than doubled to 53.5% and its context window halved to 500,000 tokens, which is why long stories lose continuity.
- There is no rollback path, since the app hides model versions and eight legacy API models were retired on May 15, 2026.
- Move character rules into Project instructions and add an explicit anti-coding line to your prompt, which recovers most of the lost quality.
- If you want persistent characters without prompt engineering every session, a purpose-built companion platform will frustrate you less than fighting this release.
