Reference

SOC / SOX Application Review

Status: Living reference · Scope: The application-review checklist an operator or auditor can walk — IT General Controls (ITGC), AI requirements, dependency & failover — and the Operational Success Management foundation underneath it.

Tags: #SOC2 · #SOX · #ITGC · #AI-middleware · #failover · #separation-of-duties


Three promises, one plane

SOC-aligned controls and security are built into the data path, not bolted on — the checklist below is walkable because the mechanisms already exist.

Companion reading — why the big-machine rebuild is the wrong-size answer: The Wrong-Size Tool walks the public escalation ladder from one ERP scoping failure to material weakness, securities litigation, and delisting. This checklist is the alternative: controls that hold after go-live.

Alignment claim, stated precisely: aligned to the SOC 2 Trust Services Criteria — design-alignment, not an attestation. SOX ITGC domains are the frame external auditors test; the evidence plane here is what your signing officers draw on.

Table stakes — at a glance

<p class="uk-heading-small uk-text-secondary">The risk is an ERP-scale failure — scoping gap to material weakness to delisting. These stakes are what keep you off that ladder.</p>

The minimum bar per domain, what enforces it, and where to watch it run — the full ladder is documented in The Wrong-Size Tool. Everything below this table is the walkable detail.

DomainTable stakesEnforced bySee it live
Access controlsInvitation-only identity for humans and agents; revocation kills live sessions, not just future loginsPersonas + separation of duties, token denylist, key ceremoniesChoose your view
Change managementAuthor ≠ approver ≠ deployer; staged realms between edit and liverelease:promote capability, deploy guards, pre-ship harnessesRelease Manager view
IT operationsEvery job idempotent and re-runnable; writes buffer through outages; cache rebuildable from truthReconcile crons, cursor gates, outage replaySession view
Program developmentSecurity in the data path from design; supply chain governedPipeline security baseline; the ERP-failure acquisition lens—
AI requirementsAgents hold their own identity and a signed, capped, expiring mandate; consent parity; refusals auditedAP2 mandates (offline-verifiable), immutable agent audit logAgent permissions
Dependency & failoverFail closed when a trusted party goes dark; verification works offline; reconcile is exactly-oncePublished keys, buffered replays, the anti-monolith legs—
Data & evidenceConsent enforced at the moment of the event; one session-keyed ledger joins order, tax, consent, engagementConsent event bus + the universal revenue ledgerSession view

Every "see it live" link opens the running console on the store — access by invitation; uninvited visitors get the full preview catalog.

ITGC 1 · Access controls

<p class="uk-heading-small uk-text-secondary">Managing who can access systems and data — provisioning, deprovisioning, periodic review.</p>

ITGC 2 · Change management

<p class="uk-heading-small uk-text-secondary">Changes tested, approved, and documented before deployment.</p>

ITGC 3 · IT operations

<p class="uk-heading-small uk-text-secondary">Backup and recovery, job scheduling, incident management.</p>

ITGC 4 · Program development / acquisition

<p class="uk-heading-small uk-text-secondary">How new systems or major changes are developed or acquired.</p>

AI requirements — ITGC applied to non-human actors

Dependency & failover

The controls, live

Not a slide — the running product. Two of the review's controls, photographed on this store — or skip the photographs and open the soul of the application yourself:

Choose your view — the persona selector: CISO/DPO, Designer, Revenue BA, QA/Compliance, Release Manager, Audience Manager, Funnel Wizard, Session view, Agent permissions, Firmware & SBOM Separation of duties as a product surface (ITGC 1): access is granted by invitation, and each stakeholder — the DPO, the analyst, QA, the release manager — gets exactly their operational view, including Agent permissions and the Firmware & SBOM harness. Open it live from the permission baseline on the How page.

Entitlement Service session view — one session showing 1P consent revoked, entitlement caps, revenue pane, and engagement suppressed with the note: consent revoked/reset for this identity, enforced from N+1, YouTube/social tracking withheld Real-time enforcement, evidenced (AI requirements · consent parity): one session, four consent-scoped panes. The consent revoke is honored — engagement tracking is withheld and says so — and the refusal itself is the audit record. This is the server-side answer to "what did the servers do?"

The cost of standing still — AI-speed data on batch-speed infrastructure

<p class="uk-heading-small uk-text-secondary">AI asks in milliseconds. Batch infrastructure answers in days.</p>

An estate without real-time server functions has exactly two options when agents arrive, and both are expensive: block the AI (the revenue and productivity cost of sitting out the platform shift) or let it act unverified (the compliance cost of an actor moving faster than your controls can watch). The remediation ladder is public record — the billion-dollar fine table, and the ERP failure walked to delisting. The alternative — consent, entitlement, and evidence enforced at the moment of the event — is not a program. It is middleware that is free to adopt.

The trust network is already breached

The case for offline-verifiable, short-lived, fail-closed authority is not theoretical. The things estates trust by default — the identity providers they delegate authentication to, and the package registries that ship their code — have each been compromised, recently and publicly.

The enterprise trust frameworks were the exploit path — Entra and Okta

<p class="uk-heading-small uk-text-secondary">Enterprises did not merely use Okta and Microsoft Entra; they made them the <strong>trust framework</strong>.</p>

The framework is the party every application defers to on the question is this session real? Both frameworks were then exploited, and in neither case did the attacker break the cryptography. They exploited the framework itself:

That is the structural lesson for every application review: a trust framework concentrates trust, and concentrated trust is a single point of failure with the largest possible blast radius. When the framework's key signs, everything opens; when the framework's support desk leaks, every customer's session walks out. An estate whose consent, entitlement, and authorization all reduce to "the IdP said so" inherits the IdP's breach at the moment it happens — and learns about it whenever a customer notices.

The controls in this checklist are the counter-shape: authority is verified per request against keys you publish, sessions are three-part and revocable now (a denylisted token dies at once, not at the next sync window), agent authority is a separately signed, capped, expiring mandate rather than an inherited session, and everything fails closed when a trusted party goes dark or goes rogue. The trust framework stays — Google sign-in is still the wristband — but it is one input to authorization, never the whole answer.

The supply chain shipped the breach — npm

All modern digital infrastructure — a teetering tower of blocks resting on a project some random person in Nebraska has been thanklessly maintaining since 2003 xkcd #2347, "Dependency" — CC BY-NC 2.5, xkcd.com/2347

The pattern across all five incidents: the trust network itself is the attack surface. Your dependencies ship to you; your identity provider's session is your session. A stolen token that lives for hours on batch-speed identity sync is a breach; the same token on real-time infrastructure dies at revocation — now.

What that lesson demands is exactly this checklist's spine: verify authority against your own published keys, offline, per request (no inherited vendor trust); tokens that are short-lived and single-use; controls that fail closed when a trusted party goes dark or goes rogue; and an SBOM discipline that can answer "do we ship the compromised version?" in minutes — the question every axios consumer had to answer on March 31, 2026, at whatever speed their inventory allowed.

Inverting risk into success

<p class="uk-heading-small uk-text-secondary">The same speed that makes AI a risk from the sidelines becomes the advantage the moment it belongs to the stakeholders who carry the consequences.</p>

The signing officer, the DPO, the analyst who must answer for what the servers did — each holds an operational view of their own, fed at the speed the question arrives: consent enforced at the moment of the event, authority verified per request against published keys, every refusal ledgered as it happens. The alternative is standing on the sidelines inside single-point-of-failure patterns that have already failed in public — one maintainer's phished inbox, one support system's session tokens — while a remediation program tries to boil an ocean of software that neither resolves the data nor solves the problem; that ladder has a documented ending. The inversion is the whole strategy: don't rebuild the estate to move at AI speed — put AI beside it as middleware that already moves at that speed, free to adopt, so the people who manage the consequences see them coming instead of reading about them in next quarter's disclosure.

The gaps this review finds in real estates

Four gaps recur in otherwise well-run estates. Each one is a join that doesn't exist — two systems that are individually healthy and jointly blind. This is what the review above surfaces, and what middleware closes without replacing either side.

SAP that doesn't handle RMA. The ERP owns the order ledger, but the return lifecycle runs somewhere else — a mailbox, a spreadsheet, a channel the ERP never sees. Refunds move money outside the system of record. Control consequence: revenue recognition and the completeness assertion break — the exact class of failure that becomes a material-weakness disclosure. The join: returns run as a tracked lifecycle whose refund lands back on the immutable order audit, keyed to the same session as the sale.

Customer service that doesn't link to WMS. Service makes promises — replacement shipped, return received, refund on the way — that the warehouse cannot confirm and service cannot verify. Promises without evidence. Control consequence: the operation attests to states it cannot prove; disputes are decided by whoever kept better notes. The join: the fulfillment event is the evidence — service reads the same stamped, ledgered events the warehouse writes.

Fraud that doesn't link to CRM. Fraud scoring sees transactions but not the person: no tenure, no consent posture, no history. Loyal customers get declined; serial abusers rotate identities beneath the threshold. Control consequence: the control exists but acts on incomplete data — precision failure in both directions, unmeasured. The join: one identity spine under every transaction, so risk decisions read the same customer record marketing and service do — and every decline is logged as a refusal, not silence.

Consent that doesn't link to CRM. Consent is captured at a banner and stays there. The customer record — and every downstream activation reading it — never learns about the revoke. Control consequence: the estate acts on data whose permission was withdrawn; the regulator's server-side question — what did the servers do after the revoke? — has no answer. This is the gap behind the fine tables. The join: consent is an event on the identity spine, enforced at the data path — a revoke gates segments, engagement, and agents from the next session forward, and the enforcement is itself in the ledger.

Four gaps, one shape: the record exists, the join doesn't. The wrong-size answer is replacing the systems that hold the records. The right-size answer is the station between them.

The foundation: Operational Success Management

ERP implementations fail for operational reasons — all seven documented causes are variants of the system only looked like it worked until operations asked it a question. The engineering community has a name for this: probably-working software.

<div class="uk-cover-container"><iframe src="https://www.youtube-nocookie.com/embed/DZpR0GojoWQ" title="The dangers of probably-working software — Damian Brady, NDC London 2026" loading="lazy" allow="clipboard-write; encrypted-media; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></div>

Damian Brady, NDC London 2026 — "The dangers of probably-working software" (privacy-enhanced player, no tracking cookies before you press play). The dogfood discipline below is the cure: software stops being "probably working" when the operator runs it on themselves and the evidence is public — this store's own consent, orders, and sessions feed the same ledger it sells. The NDC series also carries talks on permissions and identity — the same ground the AI requirements section walks.

The foundation here inverts each failure mode into a standing management discipline. That is Operational Success Management: operations is not a rollout phase that ends at cutover — it is the thing the product continuously manages, with evidence.

ERP failure modeThe OSM disciplineThe mechanism
Unrealistic goalsVerify against live operations before attestingGo-live wizard: never self-attest — open a real connection
Wrong expertiseArchitecture before implementationThe station, not another destination; human executes, agent verifies
Operational underestimationRun the product on itselfThe store's own consent, orders, and sessions feed its own ledger
Inadequate testingThe unhappy paths are the test suiteFail-closed denials, outage replays, revoked-consent paths deliberately exercised
Unverified vendor claimsEvery claim ships with its verificationPublic-key verify on certificates and mandates; dry-run previews; verify-after-deploy gates
Insufficient stakeholder communicationEvery stakeholder has an operational viewChoose-your-view personas — BA, QA, DPO, designer, release manager — separation of duties by design
Incomplete requirementsRequirements as checkable constraintsValidators and lint over one-time applies; a living gap register

Go-live is when management starts, not when the project ends. The controls above hold afterward because the same plane that runs the business produces the evidence — in real time, on every change, in systems you own.


Related: The Wrong-Size Tool · Cybersecurity for AI · Security & Compliance Posture · The dangers of probably-working software — NDC London 2026 · Print this page for the PDF edition.