v2.0.0 · local-first · offline runtime with SHA-256 evidence

Design engineering,
not style recommendation.

An Agent Skill, MCP server, and deterministic quality gate that checks whether an interface supports real tasks, exposes coherent states, uses maintainable component boundaries, remains accessible, and produces verifiable evidence.

Explore Quality Gates
scoped MCP tools
20scoped MCP tools
workflow stages
9workflow stages
evidence tiers
4evidence tiers
network calls at runtime
0network calls at runtime
git clone https://github.com/ztothez/ztothez-design-engineering.gitcd ztothez-design-engineeringnpm cinpx --no-install playwright-core install chromiumnpm run buildnpm test
The nine-stage workflow

One consolidated decision, nine bounded stages

Each stage links product task, truthful state, information hierarchy, interaction states, visual direction, semantic tokens, implementation, automated evidence, and attributable human review. Expand a stage to see the contract it validates.

Example profileresponsive-overview
Pass6Failure1Limitation1Unverified1

Status values above are illustrative example output from a sample profile run. They are not live results from your repository — run the gate locally to produce real evidence.

Capabilities

What actually gets checked

Scoped MCP access to architecture, Figma, design-system, UX-pattern, and usability guidance — backed by deterministic validators rather than opinion.

Contracts & briefs

Generation is blocked until intent is defined.

  • Versioned product design briefs that block generation when users, tasks, data behavior, recovery, assumptions, or acceptance evidence are materially undefined.
  • Design-deliverable manifests validating visual direction, semantic tokens, typography, composition, density, state, motion, charts, contrast, provenance, icons, and presentations.
  • Interface-trust contracts for data mode, connection, result origin, freshness, fallback disclosure, and provenance-preserving records.
  • Operational information-design contracts for decision metrics, evidence-backed findings, chart purpose, hierarchy, exceptional states, and scalable collections.

Repository audits

Static analysis against deterministic anti-slop rules.

  • Coupling and component-size thresholds.
  • Raw design values used instead of semantic tokens.
  • Mock production paths and placeholder interactions.
  • Missing network states and missing accessible names.

Browser verification

Evidence captured from a real Chromium session.

  • Responsive layout, clipping, overlap, focus order, contrast, and target sizing.
  • Keyboard behavior, text resizing, reflow, reduced motion, and media handling.
  • Console errors and network failures recorded as artifacts.
  • V2 contracts for data-mode disclosure, chart alternatives, masked dynamic regions, and checksum screenshot regression.

Corpus & benchmarks

Scored against maintained positive and negative cases.

  • Versioned corpus with provenance, per-dimension scoring, recommendation MRR, and explicit abstention checks.
  • Executable AegisOPS and SceneStart benchmark contracts.
  • Local-only portfolio registry with a disposable snapshot boundary for non-destructive cross-product benchmarking.
  • Anonymous comparison methodology with claim ledger and release-readiness validation.

Offline runtime & CI

Reproducible without network access.

  • Self-contained offline runtime with approved knowledge and a serialized retrieval index.
  • SHA-256 integrity evidence across production dependencies.
  • Exact provenance and dependency inventories with active-reference isolation checks.
  • GitHub Actions verification with retained fixture evidence.
Evidence model

Four evidence types, never conflated

Automated checks, AI-assisted review, attributable human review, and representative-user study each carry different weight. The system keeps them distinct.

  1. 01

    Automated evidence

    Deterministic source, contract, browser, network, and state checks. Reproducible and machine-verifiable.

  2. 02

    AI-assisted expert evidence

    Identifies likely usability risks. It does not represent observed user behavior and never substitutes for it.

  3. 03

    Human-expert evidence

    Requires an attributable reviewer. Open severity 3 and 4 findings become acceptance-criterion candidates.

  4. 04

    Representative-user evidence

    Records observed task performance with appropriate study context. The only tier that speaks for real users.

Hard rule. AI agents must never create human attestations or present generated observations as representative-user evidence. Open severity 3 and 4 heuristic findings become acceptance-criterion candidates that require review before contract integration.

Requirements

Runs on your machine, not ours

A self-contained offline runtime with approved knowledge, a serialized retrieval index, production dependencies, and SHA-256 integrity evidence.

  • Node.js 22 or newer
  • npm
  • Chromium for browser verification
  • Linux, macOS, or Windows with a Chromium executable supported by Playwright
Next step

Install from source or from the verified release archive, register the stdio entrypoint, then run your first audit and browser verification profile.