The Best Models for SillyTavern and What a Long Story Costs

What’s Changed: The models SillyTavern users run most are now cheap open-weight ones. GLM 5.2, GLM 5.3 Flash and Gemini 3.8 Flash top its usage on OpenRouter, while Claude models carry about a tenth of its top-20 volume. For most people the best pick is GLM 5.2 or Gemini 3.8 Flash, with Claude Opus saved for scenes that need its prose.

Picking among the best models for SillyTavern is the first real decision after setup, because SillyTavern has no model of its own.

Every reply comes from whatever API you connect, and that one choice sets both the prose and the bill.

OpenRouter, a service that sells hundreds of AI models through one API key, publishes which models SillyTavern traffic runs on, and the list surprised me. GLM 5.2 leads, open-weight models from Chinese labs (ones anyone can download and host) carry about two-thirds of the top 20’s tokens, and Claude sits at roughly a tenth despite its name for the best prose.

This guide ranks the options by that usage and prices a full 50-reply evening on each model, then covers the two default settings that shape what you pay.

If you are still new to the app itself, the SillyTavern basics come first.

The Best Models for SillyTavern and What a Long Story Costs

What Are the Best Models for SillyTavern Right Now?

The best models for SillyTavern right now are GLM 5.2 and Gemini 3.8 Flash for most people, DeepSeek V3.2 when other models refuse a scene, and Claude Opus when prose matters more than price.

SillyTavern OpenRouter token share by model family

That ranking follows real usage. OpenRouter’s SillyTavern app page lists the models its traffic ran on over the last 30 days, and on October 2, 2026, GLM 5.2 led with about 35.9 billion tokens.

GLM 5.2 is the model I point newcomers to first. OpenRouter’s model data lists it at $0.41 per million input tokens, so even a long story stays under a dollar, and SillyTavern traffic picks it more than any other model.

Not everyone wants to manage API keys and a token meter. If that sounds like a chore, Nectar AI runs character chat on its own plans, with no keys to configure.

RankModelSillyTavern tokens, last 30 daysList price per million tokens (input, output)
1GLM 5.235.9B$0.41, $3.99
2GLM 5.3 Flash26.0B$0.15, $0.50
3Gemini 3.8 Flash25.7B$0.75, $3.75
4DeepSeek V3.224.1B$0.28, $0.42
5DeepSeek V4.1 Flash17.2B$0.03, $0.63
6DeepSeek V4 Flash (July build)16.1B$0.01, $1.28
7DeepSeek V4 Pro (August build)14.9B$1.32, $3.96
8GLM 5.313.4B$1.40, $4.40
9Claude Opus 4.612.5B$5.00, $25.00
10DeepSeek V4 Flash (April build)12.1B$0.04, $0.08

Inside the top 20, DeepSeek models carried about 95 billion tokens, GLM about 75 billion, Google about 53 billion and Anthropic about 28 billion. List prices in the table come from OpenRouter’s model data on October 2, and the provider that serves you can charge more or less.

Put together, DeepSeek, GLM, Kimi, MiMo and MiniMax made up about 68 percent of that top-20 volume. OpenRouter’s roleplay rankings page says its charts “reflect adoption, not model quality or benchmark performance,” and price explains most of that adoption.

These counts only include requests sent through OpenRouter with SillyTavern’s name attached. Direct APIs, NanoGPT and local models never show up here, so the list shows what paying OpenRouter users choose and nothing about other setups.

Quality opinions split along different lines. In one ranking thread, a commenter lists DeepSeek V3.2 as the model to use when every other model refuses.

A Gemini 3.8 Flash thread that crowns it for roleplay also draws replies saying it refuses too often.

What Does a Long SillyTavern Story Cost on Each Model?

A long SillyTavern story costs from about 25 cents on GLM 5.3 Flash to almost $9 on Claude Opus, because every reply resends the whole chat as input.

Cost of a 50-reply SillyTavern evening by model

SillyTavern sends the full prompt with every reply: the character card, any lorebook entries and the chat history so far. A 50-reply evening at a 32,000-token prompt therefore bills about 1.6 million input tokens and only about 20,000 output tokens.

That ratio is the number I wish every pricing page led with. It means the input price decides most of your bill, and the output price barely moves it.

ModelList price per million (input, output)50 replies, caching off50 replies, Claude caching on
GLM 5.3 Flash$0.15, $0.50about $0.25n/a
DeepSeek V3.2$0.28, $0.42about $0.46n/a
GLM 5.2$0.41, $3.99about $0.74n/a
Gemini 3.8 Flash$0.75, $3.75about $1.27n/a
Kimi K3 (Moonshot AI’s endpoint)$3, $15about $5.10n/a
Claude Sonnet 5.5$2, $10about $4.42about $0.86
Claude Opus 4.6$5, $25about $8.50about $1.65
Claude Opus 5.5$4, $20about $8.84about $1.32

The math uses OpenRouter list prices on October 2, 2026, 400-token replies, and replies sent within five minutes of each other for the cached column.

The Opus 5.5 and Sonnet 5.5 rows include the extra tokens Anthropic’s newer tokenizer counts, explained in the next section.

Many providers of the cheaper models also list a lower price for repeated input on OpenRouter, so their real bills can land below these figures.

Kimi K3’s headline price of $0.354 per million input tokens comes from the cheapest of its 22 providers. The endpoint run by Kimi’s maker, Moonshot AI, charges $3 in and $15 out, which is the row used here.

For scale, one person in an OpenRouter versus NanoGPT thread spent about $8 over almost two months on cheap models. That same $8 buys less than one uncached evening on Opus.

The cheapest fix on any model is a smaller prompt. A summary and memory setup keeps old events in the story without resending every message, and halving the prompt roughly halves the input bill.

Why Can Claude Opus 5.5 Cost More Than Opus 4.6?

Claude Opus 5.5 can cost more than Opus 4.6 because it counts about 30 percent more tokens for the same text, which cancels its lower price per token unless prompt caching is on.

On paper Opus 5.5 is the cheaper model, at $4 per million input tokens against $5 for Opus 4.6. Anthropic’s pricing page also says Claude 4.7 and later models use a newer tokenizer that produces “approximately 30% more tokens for the same text.”

Run the same story through both and the discount disappears. The 50-reply evening above comes to about $8.50 on Opus 4.6 and about $8.84 on Opus 5.5 with caching off.

I don’t buy the idea that the newest Opus saves money by default. A blind roleplay benchmark posted to the subreddit found a similar gap between Opus 4.6 and Opus 5, with one benchmark run costing $0.37 on Opus 4.6 against $0.75 on Opus 5.

Once caching is on, Opus 5.5 reads cached tokens at $0.20 per million against $0.50 on Opus 4.6. That drops the cached evening to about $1.32 on Opus 5.5 and $1.65 on Opus 4.6, the one place the newer model wins on price.

How Do You Turn On Claude Caching in SillyTavern?

Claude caching in SillyTavern is off by default, and you turn it on in the config.yaml file inside your SillyTavern folder.

Version 1.19.0 ships with enableSystemPromptCache: false and cachingAtDepth: -1 in its default config file, which means no caching at all.

The request code applies those same two settings to Claude on Anthropic’s own API and to Claude models routed through OpenRouter.

Before:

claude:
  enableSystemPromptCache: false
  cachingAtDepth: -1

After:

claude:
  enableSystemPromptCache: true
  cachingAtDepth: 0

Save the file and restart SillyTavern so it picks up the change. The comment in the config says a depth of 0 “should be ideal for most use cases,” and it is the setting I would start with.

Caching only pays when the start of the prompt stays the same between replies. The config warns against it when a {{random}} macro or lorebook entries change the text before the chat history, because every change forces a fresh cache write at 1.25 times the input price.

The cache also expires five minutes after its last use. Slow writers can set extendedTTL to true for a one-hour cache, which costs twice the input price to write instead of 1.25 times.

Why Does the Same Model Feel Different From Day to Day?

The same model feels different because OpenRouter can send each request to a different provider running a different build, and SillyTavern’s defaults allow any of them.

One model name on OpenRouter is often many separate services. GLM 5.2 alone is sold by more than 30 providers there, starting at $0.1314 per million input tokens, and many of them run quantized builds.

What is quantization: Storing a model’s weights at lower precision, such as 8-bit or 4-bit, so it runs on cheaper hardware at some cost to output quality.

About half of those GLM 5.2 listings run at fp8, several use 4-bit formats, and a large share never state a precision. In the OpenRouter and NanoGPT thread, one user says a model they believe was GLM 5.3 came out “significantly worse” on NanoGPT than on OpenRouter. Another saw Gemini 3.8 blocked on NanoGPT for prompts that worked fine on OpenRouter.

I pin providers on any model I use for long chats, because a story that drifts between builds reads like it has two authors. SillyTavern 1.19.0 leaves the provider list empty, leaves the quantization list empty and ticks Allow fallback providers, per its settings code.

That combination lets OpenRouter answer with any provider at any precision. Here is the fix, in order:

  1. Open the API Connections panel (the plug icon) and choose Chat Completion with OpenRouter as the source.
  2. Pick your model, then open the Model Providers list and select one or two providers you trust for it.
  3. Untick Allow fallback providers, so a busy provider returns an error instead of quietly handing your chat to another one.
  4. In Model Quantizations, select fp8 and fp16 to rule out the most compressed builds.
  5. Send a test message and compare the reply against your last good session.
SymptomLikely causeFix
Same model writes worse than last weekOpenRouter routed you to another provider or a 4-bit buildPin one or two providers and filter quantizations
Errors start after pinning a providerThe pinned provider is busy or downAdd a second provider, or tick fallbacks back on for that session
Bill jumps after moving to Claude 4.7 or laterThe newer tokenizer counts about 30 percent more tokensTurn on caching in config.yaml, or stay on Opus 4.6
Claude bill stays high with caching onA lorebook entry or {{random}} macro changes the prompt start every turnMove changing content below the chat history, or drop the macro
A model refuses a scene it should handleThat model’s own training or the provider’s filterTry another provider, or a different model such as DeepSeek V3.2

Should You Buy Tokens From OpenRouter, NanoGPT or a Cheap Reseller?

OpenRouter is the safest default for buying tokens, NanoGPT is cheaper to top up, and a reseller that will not name its upstream providers is not worth the risk.

OpenRouter’s FAQ says it passes model prices through with no markup and charges its fee when you buy credits. Its pricing page puts that platform fee at 5.5 percent on the standard pay-as-you-go plan.

OpenRouter is also large, claiming 8 million users and 100 trillion tokens a month in TechCrunch’s May report.

NanoGPT charges no fee on deposits and no per-query minimum, according to its pricing docs. Its subscription is metered by daily and weekly input-token quotas, so long roleplay prompts, which are mostly input, eat into it fastest.

My rule with resellers is short. If a service will not say whose servers run its models, it does not get my card.

In September, a warning thread about Hapuppy described models down for days and a filtering model reading every output. CrofAI shut down the same month after it was caught routing requests for pricier models to cheaper flash models.

Prices far below what other providers charge were the warning sign in both cases. CrofAI promised refunds, but its payment provider banned the account, according to that thread.

How Do You Check a Model Before Building a Preset Around It?

Check a model’s end date and data terms on its OpenRouter page before you spend an evening on a preset, because popular models can vanish within weeks.

OpenRouter shows a “Going away” notice on models with a set end date. In early October 2026, the second-busiest model on its weekly roleplay chart was a free stealth model called Space Bunny Alpha, with a going-away date of October 5.

That model page also warned that prompts and completions for the stealth model may be retained by the provider. I keep private stories off anonymous free models for exactly that reason.

Model names can also change without any end date. DeepSeek announced in September that V4 Pro would route to V4.1 Flash, then a note on its pricing page said V4 Pro would stay, and the DeepSeek V4.1 Flash guide covers what that meant for roleplay.

I check that notice before I tune anything, since a preset built around one model rarely carries over cleanly to the next.

Which SillyTavern Model Should You Start With?

Start with GLM 5.2 for everyday roleplay, Gemini 3.8 Flash for polished prose on a budget, and Claude Opus with caching on only if you will pay about $1.30 to $1.70 an evening.

If you wantStart withWhyRough cost per 50-reply evening
The lowest billGLM 5.3 FlashSecond most used, $0.15 per million inputabout $0.25
An everyday defaultGLM 5.2Most used model in SillyTavernabout $0.74
Polished prose on a budgetGemini 3.8 FlashThird most used, but refuses moreabout $1.27
Fewer refusalsDeepSeek V3.2The usual fallback when others refuseabout $0.46
The best prose money buysClaude Opus 5.5 or 4.6Caching on, replies inside five minutesabout $1.32 to $1.65
No budget at allFree OpenRouter modelsDaily request limits apply$0

If I were setting up today with $10 of credit, I would start on GLM 5.2 with pinned providers and keep Opus for the scenes that earn it.

Readers with no budget can follow the free OpenRouter setup, which covers the free models worth trying and their daily limits.

Some people get this far and decide the whole stack is more work than the story. For them, Candy AI runs the model and the billing in one hosted app.

You give up choosing the model, and there is nothing to configure.

Quick Takeaways

  • GLM 5.2 is the most used model in SillyTavern on OpenRouter, and at list price it keeps a 50-reply evening under a dollar.
  • Cheap open-weight models carry about two-thirds of the tokens among SillyTavern’s top 20 OpenRouter models, while Claude carries about a tenth.
  • Claude 4.7 and later count about 30 percent more tokens, so Opus 5.5 only beats Opus 4.6 on price once caching is on.
  • SillyTavern ships with Claude caching off, and turning it on cuts an Opus evening from about $8.50 to under $2.
  • My default for a fresh setup is GLM 5.2 with one or two pinned providers and fp8 or better.

Leave a Reply

Your email address will not be published. Required fields are marked *