Elevated latency on EU region
Build queue p95 latency spiked from 22s to 78s for European customers. Caused by an Anthropic rate-limit incident on their end. Failover to our fallback model restored service.
Twenty-four / seven. Measured over the last 90 days. Each bar is one day.
Every incident, every postmortem, every corrective action — documented and public.
Build queue p95 latency spiked from 22s to 78s for European customers. Caused by an Anthropic rate-limit incident on their end. Failover to our fallback model restored service.
Connection pool exhaustion on the primary build database. Auto-scaling triggered, pool expanded, service restored within 12 minutes.
Webhook delivery from GitHub was failing for ~14% of repositories due to a misconfigured DNS rotation on our side. Resolved by reverting the rotation. RCA published.
A bug caused refresh tokens to expire 12 hours early, forcing some users to re-login. Hotfix deployed within 28 minutes. No data impact.
Routine PostgreSQL upgrade with read-only mode during the window. Expected zero data loss.
Stay in the loop — email, RSS, or webhook into Slack / Teams.