DeepSeek V4 Price Increase and What It Really Costs Roleplayers

What’s Changed: DeepSeek raises V4-Flash and V4-Pro API prices at 16:00 UTC on August 16, 2026, and adds peak and off-peak billing. The widely quoted 1,100% figure applies to one token type on one model at peak hours. A long roleplay session costs roughly four and a half times more at peak, and about double off-peak.

The DeepSeek V4 price increase lands at 16:00 UTC on August 16, 2026, and the number circulating in every thread is 1,100%. That figure is real.

It also describes a token type most roleplayers have never thought about once.

If you run a Janitor AI or Chub character through a DeepSeek key, your bill is not going up eleven-fold. It is going up somewhere between two and five times depending on one variable almost nobody is talking about, which is the clock on your wall.

Below is the full rate card, the arithmetic on a normal roleplay session, a time zone map showing when you are being charged double, and the one prompt setting that quietly multiplies everyone’s bill regardless of any of this.

DeepSeek V4 Price Increase for Roleplayers

What Changed With the DeepSeek V4 Price Increase

The DeepSeek V4 price increase replaces flat per-token rates with a peak and off-peak schedule starting 16:00 UTC on August 16, 2026.

Peak hours are 01:00 to 04:00 UTC and 06:00 to 10:00 UTC. Peak rates are exactly double the off-peak rates on every line.

DeepSeek V4 peak and off-peak pricing schedule

DeepSeek framed the change as a way to allocate resources more reasonably and to push developers toward scheduling work in quieter windows, as Engadget reported when the rates went public.

That is a sensible goal for a company running out of capacity. It also means the cheapest model in serious roleplay stopped being cheap in the way people got used to.

The number that matters to me is not the one in the headlines. It is the output rate, because roleplay burns output far harder than anything else, and that line went up on both models at both times of day.

Rate per 1M tokensOld flat rateNew off-peakNew peak
V4-Flash input, cached$0.0028$0.007$0.014
V4-Flash input, uncached$0.14$0.22$0.44
V4-Flash output$0.28$0.66$1.32
V4-Pro input, cached$0.003625$0.022$0.044
V4-Pro input, uncached$0.435$0.66$1.32
V4-Pro output$0.87$1.98$3.96

One thing that table kills on sight is the hope that off-peak means business as usual. Off-peak V4-Pro output is $1.98 against an old flat $0.87, so the quietest hour of the night still costs you more than double what yesterday cost.

There is no window left where you pay the old price.

Why the 1,100% Figure Is Not What You Will Pay

The 1,100% increase applies only to V4-Pro cached input tokens during peak hours, rising from $0.003625 to $0.044 per million.

That is the cheapest line on the entire rate card. It climbing eleven-fold moves your total bill far less than the output line climbing four-fold.

Cached input versus output token cost impact

Here is the arithmetic that gets skipped. Cached input cost about a third of a cent per million before the change, so multiplying it by twelve lands near four and a third cents.

Multiplying the output rate by four and a half takes you from 87 cents to $3.96. Output is where roleplay spends most of its money, which is why that smaller multiple moves your bill more than the big one does.

I don’t buy the panic, and I also don’t buy the reassurance going around that this barely affects hobbyists. Both readings pick one row of the table and ignore the rest. The honest answer needs a whole session costed out, which is the next section.

There is a second reason the cached-input line still stings even though it is small in absolute terms. Caching is the mechanism that made endless roleplay affordable in the first place, so the biggest percentage hike landing on the discount itself is a real signal about where DeepSeek wants long-context chat to go.

What Does a Long Roleplay Session Cost Now

A 100-message roleplay session on V4-Pro costs roughly 5.5 cents under the old rates, about 12 cents off-peak, and about 24.5 cents at peak.

That is a 4.5x rise at peak and a 2.2x rise off-peak, not the 12x the headline suggests.

I price a session the way people play it rather than in millions of tokens. Assume a character card, persona and system prompt totalling around 8,000 tokens of context per turn once the history fills up, replies averaging 400 tokens, and a stable prefix so most of that input bills at the cached rate.

Across 100 messages that is roughly 0.76M cached input tokens, 0.04M uncached, and 0.04M output. Run those three numbers against each column of the rate card and you get the totals above.

Session typeOld flatNew off-peakNew peak
100 messages, V4-Pro~$0.055~$0.122~$0.245
100 messages, V4-Flash~$0.018~$0.041~$0.082
Rough monthly, 30 sessions~$1.65~$3.66~$7.35

Those totals line up with real spending. Heavy roleplay on Chub ran 5 to 8 million tokens for under $2 a month before this change, and $2 to $5 monthly was the normal range. Tripling a $3 habit produces a $9 habit, which is annoying rather than ruinous.

When Are DeepSeek Peak Hours in My Time Zone

Peak windows are 01:00 to 04:00 UTC and 06:00 to 10:00 UTC, which land in the evening for the Americas and the working day for Asia.

Everything outside those seven hours bills at half the peak rate.

Convert those windows and the design becomes obvious. In China Standard Time the peak blocks are 09:00 to 12:00 and 14:00 to 18:00, which is a standard office day with the lunch break carved out as off-peak.

The two-hour gap in the middle of DeepSeek’s peak schedule is somebody’s lunch hour.

Your time zonePeak block onePeak block two
US Eastern (EDT)9:00pm to midnight2:00am to 6:00am
US Pacific (PDT)6:00pm to 9:00pm11:00pm to 3:00am
UK (BST)2:00am to 5:00am7:00am to 11:00am
India (IST)6:30am to 9:30am11:30am to 3:30pm
Australia East (AEST)11:00am to 2:00pm4:00pm to 8:00pm

My window is the awkward one, and if you are in the Americas yours probably is too. Pacific evening roleplay from 6pm to 9pm sits entirely inside peak, and Eastern users lose the 9pm to midnight block that is prime time for exactly this hobby.

The upside is that the fix costs nothing. Shifting a session an hour later on the East Coast, or an hour earlier on the West Coast, halves the rate on every token in it.

How Do I Stop My Cache From Breaking

Cached input only bills at the cheap rate when the beginning of your prompt is byte-for-byte identical between turns.

A single changing character near the top invalidates everything after it and bills the whole prompt at the uncached rate.

This is the setting I’d check before changing anything else about your setup. In live testing, moving a dynamic value like a timestamp or request ID to the top of a prompt template dropped the cache hit rate from 98.7% to 0.7%, according to DigitalOcean’s prompt caching breakdown. Nearly every turn then bills at full price.

Roleplay front-ends make this easy to trip over because several of them support time and date macros that people drop into the top of a system prompt for immersion.

Vague: a system prompt that opens with Current time: {{time}}. You are Elena, a botanist living in Lisbon...

Specific: a system prompt that opens with You are Elena, a botanist living in Lisbon... and closes with Current time: {{time}} as the final line before the user turn.

Same information, same immersion, and the stable text now sits in front where the cache can match it.

Swiping works against you the same way, because regenerating a reply steps backward through the message history and breaks prefix continuity, so a habit of swiping six times per turn costs far more than the message count suggests.

What to Do About the DeepSeek V4 Price Increase

The practical response is to move your prompt furniture around, drop to V4-Flash where quality allows, and shift sessions out of peak hours.

Those three changes together undo most of the increase without leaving DeepSeek at all.

Here is the order I work through when a model repricing lands:

  1. Move every dynamic macro, timestamp, date and random seed to the end of your system prompt, below the static character and persona text.
  2. Check your model name is deepseek-v4-flash or deepseek-v4-pro. The old deepseek-chat and deepseek-reasoner aliases were retired on July 24, 2026, and stale guides still list them.
  3. Switch everyday chat to V4-Flash and keep V4-Pro for scenes that genuinely need the stronger reasoning. Flash off-peak output at $0.66 is the cheapest paid option in this whole comparison.
  4. Move your session out of your local peak block using the table above. It is a one-hour shift for most people and it halves the rate.
  5. When a chat gets long enough that costs climb anyway, transplant it. Ask the bot to summarise the story so far, start a fresh chat with the same character, paste the summary into memory, and carry over the last two or three messages.

That last one is worth doing on a schedule rather than waiting for the bill to spike. The DeepSeek and Janitor setup walkthrough covers where these fields live if you have not touched your config since setup.

Worth noting that this is a different event from the Chutes proxy repricing earlier in the year. That one was a reseller raising its margin.

This one is the model provider raising the underlying rate, so it reaches you through every route including OpenRouter and direct keys.

Which Models Are Cheaper Than DeepSeek Now

After August 16, V4-Pro at peak costs more per output token than Kimi K2.5, while V4-Flash off-peak stays the cheapest paid option available.

The right move depends entirely on which of the two DeepSeek models you were using.

ModelInput per 1MOutput per 1M
DeepSeek V4-Flash, off-peak$0.22$0.66
DeepSeek V4-Flash, peak$0.44$1.32
Z.ai GLM 4.7$0.60$2.20
DeepSeek V4-Pro, off-peak$0.66$1.98
Nvidia Nemotron 3 Ultra$0.50$2.20
Moonshot Kimi K2.5$0.60$3.00
DeepSeek V4-Pro, peak$1.32$3.96
Moonshot Kimi K2.6$0.95$4.00
Moonshot Kimi K3$3.00$15.00

Kimi K3 is the one I’d skip on price alone for roleplay. It is a strong model, and $15 per million output tokens buys a lot of Flash sessions instead.

Nemotron 3 Ultra also runs free on OpenRouter with a 200 request per day ceiling, which is a real budget floor if you can live with the cap.

The wider field has moved since the roundup of DeepSeek alternatives was written, with GLM 4.7 and the Kimi K2 series replacing most of what people recommended a year ago.

Two things to watch if you route through OpenRouter rather than DeepSeek directly. Providers differ in whether they expose caching at all, so the cached-rate math above may not apply to whichever host you land on, and bringing your own key through some front-ends adds a 5% fee on top.

If the appeal of a token bill was never the price but the control, the honest comparison is against a flat monthly rate. Janitor+ runs $12.99 a month, which is more than the session table above suggests a moderate roleplayer spends on tokens. Janitor’s own built-in model, JLLM, also stays free at roughly 50 messages a day.

Where flat pricing genuinely wins is at the top end, and it wins on predictability rather than raw cost. Platforms like Candy AI fold the model cost into the subscription, so a long session on a bad night does not produce a surprise, and there is no proxy configuration to maintain. Nectar AI works the same way if you want a second option with a different character library.

My take is that predictability is worth paying a small premium for once you are running sessions most nights.

Below that, a tuned DeepSeek key is still the cheaper answer, and the V4-Pro tier verdict was written against the old rates and reads differently now.

Frequently Asked Questions

When does the DeepSeek price increase take effect?

16:00 UTC on August 16, 2026. Requests before that timestamp bill at the old flat rates, and everything after bills on the new peak and off-peak schedule.

Is DeepSeek V4 really 1,100% more expensive now?

No. That figure covers V4-Pro cached input tokens at peak hours only. A realistic roleplay session costs about 4.5 times more at peak and about 2.2 times more off-peak.

Which DeepSeek model is cheapest for roleplay?

V4-Flash during off-peak hours, at $0.22 per million input and $0.66 per million output. It undercuts GLM 4.7, Kimi K2.5 and Nemotron 3 Ultra on both lines.

Can I avoid the peak rate entirely?

Mostly, yes. Peak covers only seven hours a day, so shifting a session outside 01:00 to 04:00 and 06:00 to 10:00 UTC halves your rate. Off-peak is still double the old flat price.

Why did my DeepSeek costs jump before August 16?

Check your prompt for a timestamp or dynamic macro near the top. Anything that changes between turns breaks the cache and bills the full prompt at the uncached rate.

Do I need to change my model name?

If you are still on deepseek-chat or deepseek-reasoner, yes. Those aliases retired on July 24, 2026. Use deepseek-v4-flash or deepseek-v4-pro.

Quick Takeaways

  • The 1,100% figure applies to V4-Pro cached input at peak hours, which is the cheapest line on the rate card and not what drives your bill.
  • A 100-message V4-Pro session goes from about 5.5 cents to about 24.5 cents at peak, and about 12 cents off-peak.
  • Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, which hits 6pm to 9pm Pacific and 9pm to midnight Eastern.
  • Move timestamps and dynamic macros to the bottom of your system prompt before anything else. A broken cache costs more than the price increase does.
  • V4-Flash off-peak at $0.66 per million output is still the cheapest paid option, so switching models matters less than switching tiers and hours.

Leave a Reply

Your email address will not be published. Required fields are marked *