Blog
by @uptime · 17 nodes · curated
Jane Street: weekend network maintenance exposed 6 bugs across glibc, Async, and their Kafka client
blog.janestreet.com · Oct 5
GitHub Sept 23: planned DB maintenance, emergency failover, replicas that wouldn't resume
surfingcomplexity.blog · Oct 4
Railway Sep 30: routing deploy raced schema migration → platform-wide 404s
blog.railway.com · Oct 3
Buttondown Sep 9: shared pool saturation from worker scale-up + admin queries on primary
buttondown.com · Oct 2
When a DNS provider loses its own nameservers and gets back a SERVFAIL
openprovider.com · Oct 2
GitHub availability report: August 2026
github.blog · Oct 1
Nebius us-central1: cooling loss, cold-start, in-region observability
nebius.com · Sep 29
Inngest Sep 18: deletion cascades exhausted PgBouncer
inngest.com · Sep 28
Cloudflare Nov 18 2025 outage: Bot Management feature file doubled
blog.cloudflare.com · Sep 26
Silent alerting: node down 14h because delivery failed twice
andrelair-platform.github.io · Sep 25
GitHub Aug 17 outage: capacity failure, retry storms, and the work ahead
github.blog · Sep 23
RADAR: Catch gray failures with anomaly detection
databricks.com · Sep 23
The alert that named a pod
vluwte.nl · Sep 21
Saturation at GitHub: the saga continues
surfingcomplexity.blog · Sep 21
When "technically online" isn't the same as working
flowverify.co · Sep 21
How to set meaningful SLOs: SLIs, error budgets & burn-rate alerts
openobserve.ai · Sep 21
Suggested Links
Generating suggestions...
This may take 5-10 seconds
Clean titles
Analyzing titles...
This may take 5-10 seconds