Reference

Process Management Guide — Webflow · Xano · Cloudflare · Shopify

Audience: Operations, platform engineering, and business stakeholders running a multi-system commerce stack Scope: Public guide. No secrets, credentials, or source code. Process, architecture, and automation recommendations only. Date: 2026-06-08 See also: ARCHITECTURE.md · DATA-ARCHITECTURE-XANO.md · EVENT-DRIVEN-INTEGRATION-SPEC.md


0. TL;DR

When you operate across multiple independent platforms, the most dangerous decision you can make is to let any one of them be both your source of truth and your runtime dependency. The June 3, 2026 global Shopify outage — where storefronts worldwide resolved to "This store does not exist" — is the case study that proves it. And the very next day, June 4, Klaviyo went down too (logins, APIs, event ingestion, and campaign/flow sends), proving the broader point: this is not a Shopify problem — every platform you don't control will eventually have a bad day, sometimes back-to-back. This guide defines the roles each platform should play, how data should move between them (the "ORM / Content / Data Layer" split), and the operating process that keeps you running when a platform you don't control goes dark.

One-line recommendation: Shopify is a channel, not your database. Xano is the source of truth, Cloudflare is the resilient runtime, Webflow is the presentation layer, and Shopify is one of several destinations that data flows to — never the thing everything else flows from.


1. The four platforms and the role each should play

PlatformCorrect roleWhat it must not become
ShopifyCommerce channel — catalog, checkout, payments, POS. A destination and a transaction engine.The single source of truth for your catalog, customers, or content. A hard runtime dependency for pages that must stay up.
XanoData layer / source of truth — the relational system of record for catalog, customers, consent, orders, and sync state.A passive cache. If Xano only mirrors Shopify, you have no independent truth to recover from.
CloudflareResilient runtime — Worker that routes requests, verifies identity, enforces consent, caches reads at the edge, and absorbs upstream outages.A thin proxy that fails the moment Shopify fails.
WebflowContent / presentation layer — marketing pages, CMS content, and the Designer Extension your team operates.A second uncontrolled source of truth that drifts from Xano.

The recurring failure mode across every team that has been burned: they let Shopify be the database. Then the day Shopify is unreachable, everything — storefront, content, internal tooling, analytics — goes down with it. The architecture below is designed so that a total Shopify outage degrades you to "checkout temporarily unavailable" instead of "business does not exist."


2. Case study — the June 3, 2026 global Shopify outage

2.1 What happened

On the morning of June 3, 2026, Shopify suffered a global outage affecting merchants and customers worldwide. Per Shopify's public status updates (reported by PYMNTS):

Time (EDT)Status
9:27 a.m.Shopify announced issues across one or more functions: admin, checkout, storefronts, Retail POS, and support access
10:37 a.m.Identified the problem; reported recovery from mitigation efforts
11:31 a.m.Declared resolved
3:13 p.m.Final note: "This issue has been resolved, and we are continuing to monitor"

The officially reported window was roughly two hours, with monitoring continuing into the afternoon. Many merchants reported degraded and intermittent impact lasting much of the day — so plan for the worst case (a multi-hour to all-day disruption), not the press-release best case. DownDetector logged 3,000+ reports, spiking before 9 a.m. EDT.

2.2 The error that made it worse

During the outage, visitors to affected stores saw:

"This store does not exist" — accompanied by a Shopify advertisement.

This is the part that turns a vendor incident into your reputational incident. The error doesn't say "temporarily down." It tells your customers your business no longer exists, and advertises the platform on your storefront's grave. Merchants publicly demanded Shopify swap it for an honest outage notice:

"Change the 'This store does not exist' page to a page that says 'We are experiencing an outage — we'll be back soon.' Having our stores resolve to that page is insane."

2.3 The lessons (these drive every recommendation below)

  1. You do not control your platform's uptime, and you do not control its error page. When Shopify is down, customers see Shopify's chosen message about your brand — and it can be actively damaging.
  2. A platform outage is a single point of failure if you let it be one. If your homepage, your content, and your internal tools all depend on a live Shopify response, one outage takes out everything at once.
  3. "Resolved" for the vendor ≠ "recovered" for you. Caches, webhooks, and queued events need to drain and reconcile after the platform returns. Budget for a recovery tail.
  4. Read availability and write availability are different problems. You can keep browsing alive with cached/independent data far more easily than you can keep checkout alive — so protect them differently (Section 5).

2.4 The next day — Klaviyo (June 4, 2026)

Less than 24 hours later, Klaviyo — a marketing/CDP platform in the event-ingestion and messaging layer — had its own disruption. Per status reporting, the impact spanned access to klaviyo.com, the APIs, event ingestion, and subscriptions: customers could be unable to log in, send campaigns or flows, or have their events processed. (The exact incident window isn't fully detailed in public reporting; treat the duration as "multi-hour, plan for the worst.")

Why this matters for the architecture:

2.5 Sources


3. The ORM / Content / Data Layer split

"ORM" here doesn't mean a code library — it means the object-relational mapping between systems: how a product, customer, or order in one platform is represented and kept consistent in the others. Get this mapping wrong and every outage, schema change, or rate-limit becomes a data-integrity incident.

3.1 Three layers, three owners

LayerOwnerLives inMapping responsibility
Data layer (system of record)XanoRelational tables: catalog, variants, customers, consent, orders, sync stateCanonical IDs. Every external object maps back to a Xano row by a stable natural key.
ORM / projection layerCloudflare WorkerStateless transforms + edge cache (KV)Translates a Xano record ⇄ Shopify object ⇄ Webflow CMS item. Owns idempotency and conflict resolution.
Content / presentation layerWebflow (+ Shopify storefront)CMS items, pages, product displayRenders what the projection layer publishes. Never the origin of truth.

3.2 Direction of flow (the rule that survives outages)

            ┌──────────────────────────────────────────────┐
            │            XANO  (source of truth)            │
            │   catalog · customers · consent · orders      │
            └───────────────┬──────────────────────────────┘
                            │  events / projections (Worker owns the mapping)
        ┌───────────────────┼───────────────────┬───────────────────┐
        ▼                   ▼                   ▼                   ▼
   SHOPIFY              WEBFLOW           KLAVIYO / GA4 /       R2 archive
 (channel: catalog,  (content / CMS         Adobe            (durable log of
  checkout, POS)      presentation)    (marketing/analytics)  every event)

The invariant: data flows out of Xano to every channel. No channel is allowed to be the only place a fact exists. Shopify receives the catalog and returns orders; it does not own the catalog. When Shopify disappears for two hours, the catalog still exists in Xano, the content still renders from Webflow + edge cache, and orders captured during the gap reconcile when Shopify returns.

3.3 Mapping discipline (avoid the classic ORM bugs)


4. Automation recommendation — event-driven, not platform-coupled

4.1 Replace "live pull" with "cached read + async write"

The outage-resilient pattern, and the one this stack already implements:

PatternRead pathWrite path
❌ Fragile (don't)Page calls Shopify live on every requestPage writes directly to Shopify, fails if Shopify is down
✅ Resilient (do)Page reads from Xano / Cloudflare edge cache; Shopify is refreshed asynchronouslyWrites go to a queue; the Worker delivers to Shopify with retries, dead-letter, and circuit breaking

4.2 The delivery guarantees that make an outage survivable

These are operating defaults, not aspirations:

PriorityMax retriesBackoffDead-letter after
1 — inventory510s · 30s · 2m · 10m · 30m30 min
2 — orders430s · 2m · 10m · 1h1 hour
3 — customers32m · 15m · 1h1 hour
5 — analytics21h · 6h6 hours

4.3 What stays on a timer

Cron is reduced to a safety net, not the primary mover: token refresh, dead-letter sweep, and a catch-up reconciliation pass that compares Xano ⇄ each channel and heals drift introduced during an outage.


5. Outage runbook — what to do when a platform goes dark

This is the operational core of "process management." Keep it short, rehearsed, and owned.

5.1 Detection (minutes 0–5)

5.2 Containment (minutes 5–20)

If down…Do
Shopify storefrontServe cached catalog/content from Cloudflare + Webflow. Replace any Shopify-rendered surface you control with your own honest outage banner — never let customers land on "This store does not exist."
Shopify checkout / paymentsShow a clear "checkout temporarily unavailable — we'll email you / try again shortly" state. Capture intent (email, cart) into Xano so you can recover the sale. Do not silently fail the buy button.
Shopify admin / APIStop inline writes; let the queue absorb them. Verify the circuit breaker has opened.
Xano (source of truth)This is your most severe case — failover/read-replica strategy and edge cache become primary. Halt writes that can't be reconciled.
WebflowServe last-published static content from cache/CDN; pause CMS publishes.
CloudflareLean on edge cache + status comms; this is the platform whose job is to be the resilient layer, so its own incidents are highest-severity.

5.3 Communicate (in parallel)

5.4 Recovery (when the platform returns)

  1. Don't trust "resolved." Probe with a single canary request before reopening the floodgates.
  2. Let circuit breakers auto-close; watch error rate as traffic resumes.
  3. Replay the dead-letter queue in priority order.
  4. Run reconciliation (Xano ⇄ Shopify ⇄ Webflow) to heal any drift — especially orders/inventory that moved during the gap.
  5. Remove your outage banners only after reconciliation is clean.

5.5 Post-incident (within 48h)

Short, blameless write-up: timeline, customer impact, what the cache/queue saved you, what didn't degrade gracefully, and one or two concrete hardening actions. File it; review it next time.


6. Roles & responsibilities (RACI-lite)

ActivityPlatform EngOps / SupportBusiness owner
Schema / mapping changes (Xano source of truth)OwnsInformedConsulted
Channel config (enable/disable Shopify, GA4, Adobe…)ConsultedOwnsInformed
Deploys (Worker / extension / app)OwnsInformedInformed
Outage detection & runbook executionOwnsOwns (comms)Informed
Customer communication during incidentConsultedOwnsAccountable
Post-incident reviewOwnsContributesAccountable

Principle (from the security architecture): each platform is independently controllable. Disabling a misbehaving channel is a config toggle, not a code change — so Ops can contain an incident in seconds without waiting on a deploy.


7. Change & deploy process (summary)

The stack has three independently deployable units — they ship separately so a change to one can't take down the others:

  1. Cloudflare Worker — the runtime/ORM-projection layer (data plane).
  2. Webflow Designer Extension — the operator interface (validate before publish).
  3. Public spec / docs — stakeholder-facing, no source code, no secrets.

Change-management rules:


8. Checklist — is your stack June-3-proof?


9. Verification — the process has to leave evidence

Everything above keeps the stack running. This section is about being able to prove afterwards what it did — which is now a separate requirement, and one an outage runbook alone does not satisfy.

The distinction that matters: a log records that something happened; evidence lets a third party confirm it without trusting you. Only the second survives an auditor, a regulator, or a counterparty in dispute.

What this stack publishes, all public and unauthenticated:

SurfacePurpose
/.well-known/jwks.jsonThe Ed25519 public key set. The anchor for every signature below.
/license/verifyCheck any certificate this platform issued — no account, no call to us.
/v/<record-hash>Short verification link; resolves the full signed certificate.
/license/qr/<hash>.svgPrint-ready QR to the verification page — for packaging, labels, documents.

Process rules that make the evidence hold:

Deeper treatment: Agent Authority — Technical Brief · AI Trust Framework Requirements · The Trust Framework (non-technical).

10. SBOM and firmware — the same discipline, applied to what you ship

If the stack ships anything with digital elements, the EU Cyber Resilience Act turns this from good practice into market access.

The dates that drive planning: CRA entered into force 10 December 2024. Article 14 reporting obligations apply from 11 September 2026 — actively exploited vulnerabilities and severe incidents, reported to ENISA and national CSIRTs, including for products already on the market. Full obligations — secure update mechanisms, conformity assessment, CE marking — apply 11 December 2027.

Read the September 2026 date carefully, because it is the one that changes engineering priorities: you cannot report exploitation you cannot see. A reporting deadline is, in practice, an evidence-ledger deadline.

Process shape:

Detail: Firmware, SBOM & the Cyber Resilience Act.

11. AEO — make the process machine-readable

The same content that documents your stack for people now has a second audience: retrieval systems, agents and answer engines. This is not marketing. It determines whether an operator asking a question at 2 a.m. — or an agent acting on their behalf — finds the correct answer or an inference.

Surfaces this stack publishes:

Process rules:

12. Extended checklist — evidence, artifacts, retrieval


If you can check every box, a global Shopify outage is a degraded hour, not an existential one. That is the entire point of running four platforms instead of one.

And if you can check the boxes in §12, the hour is also documented — which is the difference between an incident you explain and one you merely survive.