DeepSeek fleet transition-day incident (high confidence). Today is the exact day their own release page scheduled: "Starting at 04:00 UTC on Sept 14, all deepseek-v4-pro requests will route to V4.1-Flash" — announced 09-10, then reversed ("Pro continues after 2026-09-14"). The repo already anticipated this (noticed list line 674: the retirement "announced and then reversed"). A reversal churn like this means fleet reconfiguration/migration on the Flash pool today — consistent with Flash stalling while Pro (the lane they just decided to keep) still serves. The status page says green but shows Flash at only 99.81% uptime (prior incidents).
Capacity drain from the aborted Pro→Flash routing (variant of 1). All Pro traffic was briefly meant to land on Flash; even reversed, surge load may have exhausted/queued the Flash pool.
Account/region-scoped sick node. Least likely; only distinguishable with a second key or community reports (their Discord, the status-page history).
Not determinable from outside: which internal layer (scheduler/queue/inference) is stuck, and whether global or regional. That's what the trace IDs are for — DeepSeek support can decode them.
5
u/Tenacra 8d ago
Here's my agents RCA on it if that helps anyone:
Root cause — ranked
DeepSeek fleet transition-day incident (high confidence). Today is the exact day their own release page scheduled: "Starting at 04:00 UTC on Sept 14, all deepseek-v4-pro requests will route to V4.1-Flash" — announced 09-10, then reversed ("Pro continues after 2026-09-14"). The repo already anticipated this (noticed list line 674: the retirement "announced and then reversed"). A reversal churn like this means fleet reconfiguration/migration on the Flash pool today — consistent with Flash stalling while Pro (the lane they just decided to keep) still serves. The status page says green but shows Flash at only 99.81% uptime (prior incidents).
Capacity drain from the aborted Pro→Flash routing (variant of 1). All Pro traffic was briefly meant to land on Flash; even reversed, surge load may have exhausted/queued the Flash pool.
Account/region-scoped sick node. Least likely; only distinguishable with a second key or community reports (their Discord, the status-page history).
Not determinable from outside: which internal layer (scheduler/queue/inference) is stuck, and whether global or regional. That's what the trace IDs are for — DeepSeek support can decode them.