“AI Disruption” Publication 10,500 Subscriptions 20% Discount Offer Link.
August 17 at midnight, DeepSeek’s new API pricing officially took effect.
The entire V4 lineup now uses peak/off-peak time-of-use billing. On weekdays from 9:00–12:00 and 14:00–18:00, usage is billed at peak rates; in all other periods, prices are halved. Who would have thought large models would one day also need to “shift electricity use off-peak.”
Third-party calculations show that for V4 Flash, off-peak cache-miss input prices rose about 57%, off-peak output about 136%, and peak-period output prices are close to 4.7 times the old rate.
DeepSeek is not the only one raising prices. OpenAI, Google, and Anthropic have already raised API prices or shrunk free quotas. For two years, model vendors kept driving prices down; in the second half of 2026, that price curve has started turning collectively upward.
The reason is not complicated. Growth in inference demand has outpaced growth in compute supply, and agents turn a single call into dozens of calls. When the bill changes from “one question, one answer” to “one task running for hours,” the low prices previously sustained by subsidies can no longer hold.
Capabilities keep going up — can prices still go down? That has become a required question before large models truly enter production environments.
Just now, Alibaba delivered a satisfying answer: it released Qwen3.8-Flash and simultaneously open-sourced Qwen3.8-Flash-Next (they are the same model; the former is the API product name, the latter the official open-source name).




