Umans DeepSeek V4 Flash (lab) Experimental
317.8tok/s
throughput · p50 · last 5 min
1.69s
TTFT · p50 · last 5 min
100.00%
uptime · 24h
DeepSeek V4 Flash as a Labs experiment, open for a short test window: temporary, not a permanent id. DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release, on a 1M-token context. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Access is seat-gated through the Labs page while an experiment is live. It is offered at limited capacity and availability, so expect it to be flaky and to go down under load: crash it, give it a moment, and try again. When the window ends, the model keeps serving as the pay-per-token umans-deepseek-v4-flash-0731.
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/low/high/max
Weights
Trends
Speed over the last 90 days
now 328.4 tok/s
90 days agotoday
now 1.58s
90 days agotoday
Changelog
Events for Umans DeepSeek V4 Flash (lab)
Aug 32026
The V4 Flash lab continues as umans-deepseek-v4-flash-0731-lab Testing
The V4 Flash pilot closed at the pay-per-token release: the production id now bills per token, so the seat-gated pilot on it ended rather than charge anyone by surprise. The lab reopens on the new umans-deepseek-v4-flash-0731-lab id with a smaller cohort - free, seat-gated, same experimental capacity as before. The model keeps serving as umans-deepseek-v4-flash-0731 regardless: the model stays.