Real-time and historical availability for every Crimson service. We publish uptime transparently and post a full timeline for every incident — resolved or ongoing.
Rolling availability per service, measured from external probes in five regions. Our contractual SLA is 99.9% on the Core API and Dashboard.
Each entry links to a full post-mortem. We commit to publishing root-cause analysis within five business days of any Sev-1 or Sev-2 event.
A primary Postgres node in our us-east cluster degraded under load. Automated failover to the standby replica completed in ~22 minutes, during which the Core API returned elevated p99 latency and a small percentage of 503 responses.
Post-mortem: Failover was slower than target because connection draining wasn't tuned for our peak concurrency. We've reduced the failover budget to under 90 seconds and added synthetic failover drills to our weekly game-day.
A pre-announced 35-minute maintenance window migrated the Talent Analytics warehouse to a new columnar engine. Read APIs stayed available on cached data; forecast recomputation was paused for the duration.
Post-mortem: Completed on schedule with zero data loss. The migration cut analytics query times by roughly 3× — see the changelog for details.
A backlog in the event-delivery queue caused webhook notifications to arrive up to 12 minutes late for approximately two hours. No events were dropped — all were delivered once the backlog cleared, and the retry system preserved ordering per endpoint.
Post-mortem: A misconfigured autoscaling threshold prevented the delivery workers from scaling out. We've lowered the threshold and added queue-depth alerting that pages on-call before customer impact.
No maintenance is currently scheduled. When work is planned, we announce it here and by email at least 72 hours in advance, and we schedule windows for 02:00–05:00 UTC Sundays to minimize customer impact. Enterprise customers with strict change-freeze requirements can opt into extended notice.
Subscribe to receive real-time email notifications for incidents, maintenance windows, and post-mortems across every Crimson service.
Subscribe to updates