03 · The operating system

Nine stages.
One quality conversation.

Every stage connects a purpose, AI contribution, human responsibility, artefact, checkpoint, quality gate, risk, and measure.

01

Locate the work

Choose the stage that describes the decision being made now.

02

Read the boundary

See what AI can accelerate and what the designer must own.

03

Run the checkpoint

Apply the quality gate before the work moves forward.

Quality travels with the work.

The guardrail layer connects intent, evidence, trade-offs, accountability, and monitoring across all nine stages.

01Goal
02Risk level
03Quality dimensions
04Evidence
05Trade-offs
06Impact
07Human sign-off
08Monitoring

Stop / rework Unacceptable risk, weak evidence, or unresolved critical failure.

Quality drift Re-enter when behaviour, content, models, or outcomes change.

Outcome

Connect design to measurable user and business change.

Usability

Make tasks understandable, efficient, and recoverable.

Inclusion

Support abilities, contexts, languages, devices, and access needs.

Clarity

Make purpose, choices, status, and consequences visible.

Coherence

Align states, channels, content, components, and agents.

Trust

Communicate capability, limitations, uncertainty, and provenance.

Control

Enable edit, pause, reject, undo, exit, and recovery.

Safety

Reduce privacy, security, bias, manipulation, and misuse risk.

Resilience

Design for errors, latency, interruption, and model failure.

Choose where you are.

The process is a continuous loop. Stage 09 feeds validated learning into the next Stage 01 charter.

01
Formerly Requirements

Opportunity Framing

Establish what problem is actually being solved, for whom, and under what constraints, before any research spend begins

+
Human owns

Facilitation of the live workshop; reading organizational politics and unstated authority; all commercial/contractual decisions; final charter sign-off

AI accelerates

Structured intake at scale (parallel async interviews instead of serial scheduled ones); first-pass synthesis and conflict-flagging so the workshop starts at decision-making, not information-gathering

Inputs
Prior RFP/SOW/contract language, existing product documentation, stakeholder roster, past project retrospectives, any existing analytics or support-ticket access
Deliverables
Stakeholder interview synthesis, conflict log, draft-then-approved project charter (objectives, scope boundaries, constraints, success metrics)
Human checkpoint
(1) Workshop output reviewed against AI-drafted brief, not rubber-stamped. (2) Charter formally signed off by accountable human before Stage 02 begins, this is the Framework's hardest gate; nothing downstream is trustworthy if this is skipped
Gate to advance
Charter has a named accountable approver, explicit success metrics, and every flagged stakeholder conflict has a documented resolution (not just a note that it exists)
Risks
Treating the AI-drafted brief as final instead of a workshop starting point; AI missing context that only exists in someone's head and was never written down or said in an interview; stakeholder data itself being sensitive in regulated verticals
Measures
Time from kickoff to signed charter; number of stakeholder conflicts surfaced pre-workshop vs. discovered mid-project (lower post-charter discovery = better framing); charter revision rate during Stage 02–03
Enterprise outputs
Signed project charter, the artifact that appears in the SOW and gets referenced in every later stakeholder conversation about scope
02
Formerly Research

Intelligence Gathering

Build a verified fact base, user behavior, market landscape, competitive structure, grounded in the charter's success metrics from Stage 01

+
Human owns

All live ethnographic research (AI cannot read micro-expressions, tone, or cultural subtext); persona/JTBD validation against real interviews before anything is trusted as fact; ethical data curation; strategic differentiation calls

AI accelerates

High-volume ingestion and clustering that would take a human team weeks; first-draft persona/JTBD synthesis for human validation; competitive feature mapping at a breadth no human team matches in the same time

Inputs
Signed charter, support tickets, app store/NPS reviews, product analytics access, competitor product list, prior research repository
Deliverables
Thematic insight clusters (with outliers flagged, not discarded), draft personas + JTBD statements, competitor feature/interaction matrix, research repository entries
Human checkpoint
Every AI-drafted persona is validated against real research before use, never presented to a client as fact until validated. Ethnographic research is scoped and conducted by humans, full stop
Gate to advance
At least one persona/JTBD statement per primary user segment has been validated against real interview data (not just AI-synthesized); competitor matrix reviewed for strategic relevance, not just completeness
Risks
Treating AI-generated personas as ground truth without validation, the single most-cited 2026 failure mode in this space; AI cannot observe behavior or feel frustration, only synthesize what's already been said. Outlier-discarding if the agent isn't explicitly instructed to flag rather than drop anomalies
Measures
% of personas validated against real research before use; breadth of competitive coverage vs. time spent; ratio of AI-surfaced themes that survive human review unchanged vs. corrected
Enterprise outputs
Research repository, validated persona set, competitor landscape deck, the artifacts that ground every later stage's decisions and that clients expect to see in a discovery readout
03
Formerly Synthesis / Problem Framing

Insight Generation

Convert the Stage 02 fact base into a defined, ownable problem statement. Promoted to its own stage because problem framing is the highest-leverage human-judgment step in the process and was previously buried as a sub-bullet

+
Human owns

Final authorship and ownership of the problem statement, non-delegable. Judging which candidate framing captures real user empathy vs. which merely reads well

AI accelerates

Generating a wide candidate set of problem framings so humans are choosing from options, not starting from a blank page; checking each candidate framing against the evidence base for support

Inputs
Validated personas/JTBD, thematic insight clusters, competitor matrix, signed charter's success metrics
Deliverables
Problem statement (human-authored), 3–5 candidate HMW statements with the selected one marked, evidence trace linking the statement back to Stage 02 data
Human checkpoint
The final problem statement must be authored, not merely selected and left verbatim, by a human. Verbatim adoption of an AI-drafted statement without editorial ownership is a checkpoint failure, not a shortcut
Gate to advance
Problem statement is specific enough to be falsifiable (a bad concept in Stage 04 should be visibly bad against it), and traceable to at least two independent evidence sources from Stage 02
Risks
Selecting the best-worded AI candidate instead of the best-fit one, named explicitly in the Foundation doc as this stage's complacency risk; false precision (a sharp-sounding statement that isn't actually falsifiable)
Measures
% of Stage 04 concepts that can be scored as clearly on/off problem statement (a vague statement produces ambiguous scoring); stakeholder recall of the problem statement weeks later (a sign it's memorable and owned, not just filed)
Enterprise outputs
The problem statement itself, this becomes the north-star artifact quoted in every subsequent stakeholder review, and the yardstick Stage 07 validation is measured against
04
Formerly Ideation

Concept Exploration

Generate a high volume of structural and conceptual directions against the Stage 03 problem statement, then narrow to a defensible shortlist

+
Human owns

Final concept selection, taste, brand alignment, emotional intuition are non-delegable; ethical veto power, especially in sensitive contexts (healthcare, financial hardship, etc.); making the genuine innovation leaps AI can only interpolate toward

AI accelerates

Producing volume and structural diversity beyond what a human team generates in the same time; applying frameworks consistently as ideation prompts rather than replacing human creative judgment

Inputs
Approved problem statement + HMW, validated personas, competitor matrix (for differentiation, not imitation)
Deliverables
Concept set (volume, sketched), shortlist with selection rationale, anti-solutions log (what was deliberately rejected and why)
Human checkpoint
Concept selection is a human decision recorded with rationale, not just a pick; any concept touching a sensitive context gets an explicit ethical review before shortlisting
Gate to advance
Shortlist has documented rationale traceable to the Stage 03 problem statement, and at least one anti-solution was considered and explicitly rejected (proof the exploration was genuinely divergent, not a single idea dressed multiple ways)
Risks
Mistaking volume for progress; shortlisting the most-generated pattern because it appeared often, not because it's right; AI's structural interpolation reads as "safe" and can quietly narrow the option space toward convention
Measures
Volume-to-signal ratio (# concepts shortlisted / # generated) tracked over time, a flat or declining ratio means the agent isn't learning the team's taste and prompting needs recalibration; time from problem statement to shortlisted concept
Enterprise outputs
Concept board with selection rationale, the artifact presented in a concept-review stakeholder session
05
Formerly Information Architecture

Experience Architecture

Structure the shortlisted concept into navigable, cognitively sound information hierarchy

+
Human owns

Cognitive-load validation by feel, not just by metric; localization and cultural adaptation; the specific discipline of removing steps, AI tends to map existing logic faithfully rather than challenge it

AI accelerates

Fast first-draft structure generation from standard mental models; consistent labeling suggestions; quantifiable load simulation as a starting diagnostic, not a verdict

Inputs
Shortlisted concept, personas/JTBD, any existing taxonomy/content inventory
Deliverables
Sitemap/nav tree, taxonomy/labeling doc, cognitive-load simulation report, draft microcopy
Human checkpoint
A structure that scores well on AI's click-count metric still gets a human "does this actually feel simple" pass before sign-off
Gate to advance
At least one full simplification pass has removed structure the AI proposed as-is (evidence the human step wasn't skipped); taxonomy reviewed for the target market's language and cultural context, not just English-default output
Risks
A structure that's technically efficient (few clicks) but cognitively unintuitive; AI reinforcing existing complexity because it maps reality rather than questioning it; taxonomy that reads correctly in the source language but breaks in localization
Measures
Clicks-to-goal for top user tasks; rework cycles between IA and later stages; % of AI-proposed structure that survives the human simplification pass unchanged (a high % may mean the pass isn't rigorous, not that the AI got it right)
Enterprise outputs
Approved sitemap and taxonomy doc, referenced directly in the dev handoff package at Stage 08
06
Formerly Interaction & Visual Design

Experience Crafting

Apply behavior, motion, and visual language to the architected structure, the concept becomes a specific, branded, functioning experience

+
Human owns

Bespoke motion choreography, the specific arc/timing/easing that creates brand feel, which AI cannot originate; edge-case design, which requires empathy for user frustration AI doesn't have; cultural sensitivity review of generated imagery; final taste call on visual direction

AI accelerates

Mechanical application of design-system rules at scale and with near-100% consistency; rapid generation of styling variations for human curation; standard interaction pattern implementation, freeing human time for the bespoke moments that matter

Inputs
Approved IA, brand/design system, interaction requirements from the concept
Deliverables
Design-system-compliant high-fidelity screens, functional interactive prototype, documented edge-case states, motion spec for signature interactions
Human checkpoint
Every generated carousel/auto-advancing/timed interaction gets an accessibility check before it reaches Stage 07 testing, this is where several documented 2026 accessibility complaints originate. All imagery/iconography clears cultural-sensitivity review before client presentation
Gate to advance
Design-system token/component compliance verified (not manually eyeballed); every primary flow has a defined error, empty, and offline state, not just the happy path; at least one signature interaction has human-authored choreography, not AI-default timing
Risks
AI-default micro-interactions creating accessibility friction in contexts the agent wasn't told to consider; visual variation volume creating decision fatigue instead of clarity; cultural imagery risk in generated illustrative assets, especially for multi-market work
Measures
Design-system compliance rate of AI-generated screens (target near-100%); # of edge-case states defined per flow; accessibility issues caught at this stage vs. caught later at Stage 07 (earlier catch = cheaper fix)
Enterprise outputs
High-fidelity, interactive, design-system-compliant prototype, the artifact used in stakeholder demos and Stage 07 usability sessions
07
Formerly Testing/Prototyping

Experience Validation

Validate the crafted experience against real users and objective heuristics before build commitment

+
Human owns

Live moderation where reading body language, tone, and hesitation matters; root-cause interpretation of failure (the "why," not the "that"); final call on what gets fixed, balancing AI-reported heuristics against real constraints

AI accelerates

Mechanical, exhaustive, and fast heuristic/accessibility scanning; pre-interview volume via AI-moderated or synthetic sessions to focus scarce human moderation time on the sessions that need it most

Inputs
Interactive prototype, defined edge-case states, problem statement (as the validation yardstick)
Deliverables
WCAG audit report, heuristic findings log, usability session synthesis (human-moderated + AI-moderated combined), prioritized fix list
Human checkpoint
Any high-stakes test (regulated-industry flows, first-time-user critical paths) is human-moderated, not delegated to synthetic/AI-moderated sessions alone. Findings prioritization is a human call, explicitly weighing business/technical tradeoffs the agent doesn't have visibility into
Gate to advance
Zero unaddressed WCAG failures at the level committed to in the charter; every "user failed task" finding has a documented root-cause interpretation, not just a fail count; fix-priority list has explicit human sign-off
Risks
Over-trusting synthetic user testing as a proxy for real user reaction, explicitly flagged in current 2026 UX research practice as the most common misuse; treating a clean heuristic scorecard as proof of a good experience rather than proof of absence of *known* mechanical issues
Measures
% of accessibility/heuristic issues caught pre-human-testing (should trend toward "all the mechanical ones"); task success rate and time-on-task from human-moderated sessions; ratio of AI-flagged vs. human-discovered findings (a healthy practice finds real issues in both lanes, not just one)
Enterprise outputs
Usability test report and remediation plan, often the artifact a regulated-industry client's compliance/legal stakeholders review directly
08
Formerly Handoff/Delivery

Build Enablement

Translate the validated design into developer-ready code, specs, and assets without losing design intent

+
Human owns

Pixel/motion-nuance QA, AI translation reliably misses spacing and timing nuance a trained eye catches; technical negotiation between design intent and backend reality; security/privacy review, applied with *more* scrutiny to AI-generated code, not less

AI accelerates

Fast, consistent code generation and spec-tagging at a volume that would otherwise consume significant engineering time; asset optimization and multi-resolution export

Inputs
Validated prototype, remediation-complete design, design system
Deliverables
Semantic code (HTML/CSS, SwiftUI, Jetpack Compose as applicable), tagged design spec, optimized asset package, component unit tests
Human checkpoint
Every generated-code artifact passes the same security/privacy review as human-written code, explicitly never skipped because "the AI wrote it." Visual QA against design intent before dev merge
Gate to advance
Security/privacy review signed off; visual QA confirms spacing/motion/timing match design intent within defined tolerance; any design-backend contradiction has a documented resolution, not a silent workaround
Risks
Complacency, the single named risk for this stage in the Framework: reviewers who skip scrutiny of code specifically because it looks clean and AI-generated code often does look clean; silent design-backend contradictions that get "resolved" by an undocumented workaround instead of a real negotiation
Measures
Design-QA rework cycles between generated code and final dev handoff; time from validated design to dev-ready spec package; security review findings rate on AI-generated vs. human-written code (should converge, not diverge, over time)
Enterprise outputs
Dev-ready spec + code + asset package, the artifact that closes the design phase of the SOW and opens the engineering handoff
09
Formerly Analytics

Continuous Learning

Monitor the shipped experience continuously and feed findings back into Stage 01 of the next cycle. Decoupled from Stage 07 because this is an always-on loop, not a one-time gate

+
Human owns

Root-cause investigation, an anomaly may be an external event, not a design flaw, and only a human has the context to tell the difference; the Engagement Trap guardrail, explicitly preventing AI from optimizing toward addictive patterns just because the data says they "work"; final shipping decisions, balancing data against brand promise and long-term trust

AI accelerates

Continuous, tireless monitoring at a scale no human team sustains manually; statistical rigor on significance testing; first-pass anomaly flagging that focuses human attention

Inputs
Live product telemetry, A/B test configuration, shipped design's success metrics (from the original Stage 01 charter)
Deliverables
A/B test results with significance reporting, anomaly log with root-cause notes, quarterly (or cadence-appropriate) learnings digest feeding the next Stage 01
Human checkpoint
No variant ships to 100% without a human decision. Every flagged anomaly gets a human-authored root-cause note before it's closed, not just an "acknowledged" status
Gate to advance
Risks
The Engagement Trap, optimizing for a metric that technically improves while harming user trust or well-being (e.g., infinite scroll, dark-pattern-adjacent nudges); anomaly-flagging blind spots, the agent flags what it's tuned to notice, "no anomaly flagged" is not the same as "nothing to review."
Measures
Time from anomaly detection to human root-cause triage; % of quarterly learnings that actually get written into the next Stage 01 charter (the real test of whether this is a loop or a report); shipped-variant business outcome vs. predicted outcome (calibration check on the whole framework)
Enterprise outputs
Live performance dashboard + quarterly learnings digest, the artifact that proves the Framework compounds over cycles instead of resetting to zero each project