Trouble Accessing SDP

Incident Report for Semarchy Data Platform

Postmortem

Summary

On 9/8/2026, starting around 15:43 UTC and lasting approximately 29 minutes, some users were unable to connect to Semarchy SDP tenants in both the EU and US (NA) production regions due to a DNS resolution issue. The issue was tied to a Cloudflare zone configuration change and was resolved by rolling back that change.

Impact

Some users could not connect to tenants across both EU and US production regions due to intermittent DNS resolution failures.

Timeline (all times UTC, 9/8/2026)

  • 15:43 — Connectivity issue reported for an EU tenant.
  • 15:45 — Issue confirmed for tenants in the US region as well.
  • 15:46 — Investigation begins.
  • 15:47–15:50 — Connectivity confirmed inconsistent across tenants; DNS suspected.
  • 15:50 — Recent Cloudflare zone changes checked.
  • 15:52 — Cloudflare suspected as the cause.
  • 16:03 — Cloudflare change identified; rollback started; connectivity begins recovering.
  • 16:04–16:06 — Mixed connectivity during rollback propagation.
  • 16:12 — Connectivity fully restored.

Root Cause

An overscoped process for creating a new Cloudflare zone — intended only to provision the zone for later configuration — caused Cloudflare to begin migrating nameserver (NS) records as part of its standard zone setup, before the zone was actually ready to be used. That NS record migration is what caused the DNS resolution failures for existing tenants. Reverting the NS record migration restored normal DNS resolution.

What Went Well

  • The issue was flagged quickly and consolidated in one thread rather than opening separate reports.
  • The symptoms were connected to a recent Cloudflare change within ~7 minutes of investigation starting.
  • Rollback resolved the issue with no reported data loss.

What Didn't Go Well

  • No automated alert fired — the incident was first detected by an end user, not internal monitoring.
  • Impact was inconsistent across users, which slowed shared understanding of scope.
  • Rollback propagation took a few minutes, producing conflicting "fixed / still broken" reports.
  • The zone-creation process was broader in scope than intended, triggering NS record migration when only provisioning was needed.

Metrics

  • Time to Detect (via customer/internal report): ~0 min (15:43 report; no monitoring alert fired)
  • Time to Mitigate (rollback started): ~20 min (16:03)
  • Time to Resolve (full recovery confirmed): ~29 min (16:12)

Follow-up Actions

  • Update the Cloudflare zone creation runbook to clarify that when provisioning a new zone for later configuration, it's expected — and correct — for Cloudflare to flag the zone as incomplete or pending until it's intentionally activated. That warning should be left alone rather than acted on, since resolving it prematurely triggers the NS record migration that caused this incident.
  • Add DNS resolution monitoring from multiple external vantage points. Our internal monitoring only checks resolution from one network path, which is why it didn't fire during this incident.
Posted Sep 09, 2026 - 13:02 UTC

Resolved

This incident has been resolved.
Posted Sep 08, 2026 - 16:36 UTC

Monitoring

A fix has been implemented and we are monitoring the results.
Posted Sep 08, 2026 - 16:04 UTC

Investigating

Some users are having trouble accessing their tenants
Posted Sep 08, 2026 - 15:56 UTC
This incident affected: North America and EMEA.