Commercial
Shaping the deal and the market story so the engagement is priced honestly against confidence and positioned to win.
Domains are types of work. They do not change and they do not have levels. Capabilities are the named outcomes we promise inside a domain.
How a capability is executed is a separate scale — L1 Guided Execution, L2 Practitioner, L3 Advanced Lead, and L4 Capability Ownership. That scale lives with capabilities, not with domains.
Shaping the deal and the market story so the engagement is priced honestly against confidence and positioned to win.
Understanding the problem well enough to make the go, redirect, or stop call.
Making the thing real — from the smallest proof to durable, ownable production.
Making it demonstrably true that the work does what it was supposed to — confidence from what the client can see, not what they're told.
Making the client's people more capable than they were.
Making the system more capable than it was.
| Level | Description |
|---|---|
| L1 Guided Execution | Operates safely using defined guardrails and templates. Non-specialists can execute without inventing method. |
| L2 Practitioner | Executes end-to-end independently and handles edge cases without escalating the method. |
| L3 Advanced Lead | Builds guardrails, handles extreme complexity, and defines the standards others execute against. |
| L4 Capability Ownership | Agency-wide accountability for capability maturity, L1 guardrail creation, and quality. Distinct from L3 execution. An L4 is responsible for the capability existing and remaining fit for use. |
Frames the deal as an envelope you can commit to: clear boundaries, room for what discovery turns up, so scope stops being a running fight.
A commercial envelope with fixed boundaries and batched trade-off checkpoints instead of line-item disputes.
Define scope by outcomes + non-negotiable boundaries, not a feature list; build an explicit emergence allowance into the budget; set checkpoint gates where trade-offs are decided together.
No L1Deciding which boundaries are truly non-negotiable and how much emergence to fund is a commercial judgment a template can frame but not make.
An estimate that says plainly what we're sure about and what's still a bet, so you're never paying for false precision.
An estimate that separates proven cost from high-uncertainty bets, with confidence shown per line.
Pull analog baselines from delivery history; tag each line proven / bounded / bet; show confidence bands, never a blended number.
No L1Tagging a line proven / bounded / bet, and setting the band width, is a judgment about how much uncertainty the evidence actually leaves — structured by a guardrail, decided by a person.
None assessed.
A price that stays fair when the full shape of the work can't be known yet: fixed where it's proven, staged where it isn't.
A staged price — fixed rates for known work, investment staged as clarity grows.
Stage price across confidence gates; itemize the emergence allowance in the agreement; tie phases to signal thresholds from delivery.
No L1What to fix, what to stage, and how to price the emergence allowance is a risk call between firm and client, not a template output.
None assessed.
Finds the walls that can't move (legal, technical, contractual) before the plan is built, so they're inputs and not late surprises.
A constraint map the plan is built against, not discovered late.
Interview across four lanes (legal, regulatory, technical, operational) on a fixed script; sort constraint vs preference on the rubric; publish the set before sequencing.
Four-lane interview scriptconstraint-vs-preference rubricconstraint-set record templateL1 runs the script and populates the rubric; arbitrating ambiguous constraints is L2.
None assessed.
A clear call on whether the thing is worth building and what to do first, made in the room with the reasoning behind it.
A proceed / redirect / stop decision and a defensible "what first."
Score candidates against the problem sentence, constraint set, and risk shortlist; sequence only what survives; record the call; treat stop as a valid outcome.
No L1The proceed / redirect / stop call is the judgment; a guardrail can assemble the inputs but can't own the decision or its consequences.
Agreement on what winning actually looks like, in numbers, before anyone designs or builds.
Up to three agreed, measurable targets locked before work starts.
Translate the lever into ≤3 measurable targets in-session; run a kill-check that each is measurable; hand instrumentation to Signal design.
Target template (lever → ≤3 targets)measurement kill-checkSignal design handoff stubL1 runs the template; judging whether a complex target passes the kill-check is L2.
None assessed.
Cuts through the ask to the one thing the work really has to solve — a single sentence everyone can line up behind.
A single-sentence problem statement that aligns decisions across teams.
Run a framing session with sponsor + a blocker; ladder symptom → lever live; draft, swap-test, and lock the sentence in the room.
Problem-sentence templatesymptom→lever laddering prompt packswap-test checklistone worked exampleL1 facilitates from the prompt pack and drafts the sentence; resolving a sentence that fails the swap test is L2.
None assessed.
Separates the few risks that could sink the whole thing from the noise, so attention goes where it actually matters.
A shortlist of collapse-level risks; the rest stays off the steering agenda.
Run an assumption dump + failure premortem; sort by collapse potential (impact × irreversibility); keep only what changes the proceed call or first slice.
Dump + premortem prompt packcollapse-potential sort rubricrisk-shortlist templateL1 facilitates the dump and applies the score; judging irreversibility ratings is L2.
None assessed.
Gets the people who could stop the work to agree, up front, on what good looks like — in language that outlasts us.
A shared decision standard locked in executive language that holds across org shifts.
Map who can block or bless; put the lever in front of them in a working session; secure sign-off on a standard that holds without us.
No L1Reading who can actually kill the work and moving them to a shared standard is live political judgment; the kill-map is prep, the room is not L1.
None assessed.
AI that actually works on your real problem and your real data, not a demo that falls over in production.
A production-ready AI system grounded in your data, with accuracy and limits you can verify.
Stand up an eval harness with golden sets before architecture; ingest domain data into model-agnostic context pipelines; gate deploy on eval scores, not demo quality.
Eval-harness scaffoldcorpus-ingestion checklistreliability-benchmark templateL1 runs scaffolds and eval suites; modifying context pipelines is L2.
Builds and modifies the context pipelines; tunes reliability against the eval bar independently.
Sets the eval methodology and reliability architecture others build to; owns what "trustworthy enough to ship" means.
None assessed.
The system underneath, built to fit your stack and hold together — something your own team can extend, not just inherit.
A clean architecture matched to your stack that internal teams can extend.
Choose proven patterns matched to the constraint set; build clean seams across services, data, and APIs; document interfaces well enough for a third-party dev.
Approved architecture-pattern catalogueseam-documentation templateconstraint-citation formatL1 implements established patterns; architecting custom seams is L2.
None assessed.
A codebase and release path clean enough that your team keeps shipping long after we've gone.
An automated release pipeline and readable repo ready for handoff.
Configure CI/CD in the client's infrastructure; hold the codebase to readability and modular-interface standards; test against the "next change without the original author" bar.
Release runbook templatecodebase-readability checklistCI/CD scaffoldL1 runs CI/CD scaffolds and formatting rules; changing core pipeline logic is L2.
None assessed.
The part users actually see and touch, built for real: a working product, not a prototype.
A working, usable product shipped to live preview, backed by a reusable component library.
Assemble from design tokens + a Figma/code component library; ship increments to preview; check token use and accessibility on every build.
Component token sheetpage assembly template (Figma + code scaffold)in-token component-generation prompt packL1 assembles from tokenized components; defining new tokens or state hooks is L2.
Owns state and edge-case interface behaviour without escalating the method.
Sets the tokens, the component contracts, and the motion language others build to.
Makes sure it holds up under real traffic, real users, and real pressure, not just in the demo.
A resilient, observable system verified under load and real failure.
Run security, load, and observability against the constraint set + risk shortlist; wire monitoring/tracing into the deploy; map each fix to a named risk.
Hardening checklist (scoped from constraints + risks)observability wiring scaffoldrisk-remediation report templateL1 deploys scaffolds and runs hardening scripts; diagnosing complex load failures is L2.
None assessed.
The smallest real, working thing that answers a live question — proof over promises, early and cheap.
A running artifact plus signal, delivered early.
Frame one assumption per slice; build only what tests it in Cursor / Claude Code, deployed to a preview; log the decision (continue / pivot / stop) against the signal.
Slice charter (one assumption + success/fail + kill date)preview deploy scaffoldone-assumption gateL1 wires the scaffold for a defined slice; refactoring logic across slices is L2.
None assessed.
Clear, evidence-backed sign-off that the work does what was agreed, not a subjective "looks done to me."
Sign-off backed by pass/fail proof against criteria set up front.
Validate against the outcome targets and eval bar; record pass/fail per criterion; require verification before "complete."
Acceptance checklist (wired to agreed criteria)pass/fail record templatecriteria-verification guideL1 runs the checklist against defined cases; ruling on edge-case exceptions is L2.
None assessed.
Progress you can see running, with the numbers behind it — not a status deck that tells you it's fine.
Interactive review of running software alongside its live signal.
Demo only from live preview; walk the signal alongside the functionality; capture decisions in the session.
Evidence-review run sheetlive-preview default dashboardfeedback-capture logL1 runs the walk-through from the run sheet; handling off-script technical questions is L2.
None assessed.
Deciding up front what would actually count as proof, and building in the means to measure it.
The measures and thresholds that make results evidence, set before build.
Define metric, source, and threshold before slice build; embed measurement hooks in components/services; record the bar so results can't be re-litigated.
Signal spec template (measure / source / success / failure / threshold)instrumentation starter kitsignal-vs-demo gateL1 embeds standard hooks from the kit; defining custom metric formulas is L2.
None assessed.
You can watch the work being true as it's built, any time, instead of waiting for a weekly report.
Continuous access to live preview and status, not status packs.
Deploy a preview per active branch; expose live health/progress to stakeholders; replace at least one status meeting with the live link.
Preview-deploy scaffoldlive status-view template"no status pack where a live view exists" ruleL1 configures preview links and status views; fixing a broken delivery channel is L2.
None assessed.
Puts the work through its paces, including the ugly edge cases a happy-path demo never hits, and shows what held and what didn't.
Coverage that proves it holds, with performance boundaries and known failure modes named.
Run automated suites + generated edge cases; for AI features, run eval suites against the signal bar; archive a summary naming passes and misses.
Test-plan templateautomated eval-suite scaffoldfailure-summary format (must name misses)L1 triggers runs and fills the failure report; authoring custom generative suites is L2.
None assessed.
Your people in the room for the hard calls, so they can run and extend the system themselves — not just own a repo they didn't help shape.
A team that can maintain and extend the system on its own.
Pair client people on the calls that set the method; transfer the reasoning, not just the docs; record decision rationale for later reference.
No L1You have to be in the hard calls to transfer the judgment behind them; there's no guardrailed way to teach judgment you don't yet hold.
None assessed.
Getting the wider organization to actually trust and use what was built, so it gets used instead of shelved.
Measurable adoption across the people who didn't build it.
Map affected users outside the build team; design adoption paths against their incentives and fears; track usage and treat low usage as a signal to act on.
No L1Designing an adoption path against a specific org's incentives and fears is bespoke change judgment; usage can be instrumented, the design can't be templated.
None assessed.
Leaves your people genuinely better at their craft, not just trained on this one build.
Higher capability across your engineering and product people.
Coach 1:1 against live problems in their domain; assess against a craft rubric; leave a per-person growth roadmap.
No L1Growing a person's craft against their live problems is individual judgment; a rubric can assess, the coaching itself isn't guardrail-executable.
None assessed.
The controls, review, and stop-rules that let you run an AI system safely once it's yours — not just own something you can't supervise.
Explicit kill switches, guardrails, and escalation paths for AI in production.
Define operating boundaries, drift measures, and escalation triggers post-launch; build kill rules and override paths for client operators; deliver a runbook tied to known failure modes.
Governance runbook templatekill-rule spec sheetdrift-monitoring formatL1 fills the runbook from templates; defining custom kill rules is L2.
None assessed.
A clear picture of who owns what and how decisions get made once we're gone: a machine your team knows how to run.
An operating model defining responsibilities, ownership, and cadences.
Define post-launch roles, cadences, and decision rights; write it in the client's own org language; map when to execute, escalate, or change the system.
No L1Deciding who owns what and which calls can't wait is organizational-design judgment specific to each client; a template captures the format, not the decision.
None assessed.
No surprises after go-live: exactly what we still hold, what you hold, and when it changes hands.
A handoff schedule with dates, support coverage, and ownership transfers.
Draft one agreement covering warranty, support, and backlog transfer; specify exact transfer dates and boundaries; run a backlog review to confirm ownership.
Transition/warranty templatebacklog-handoff formatownership-transfer checklistL1 completes the standard checklist; negotiating custom warranty terms is L2.
None assessed.
| Agent skill |
|---|
accessible-component-audit
Audit interface components against accessibility criteria so later screens inherit a passable baseline instead of shipping inaccessible patterns.
|
engagement-change-runner
An executable pass over the source-of-truth model that answers "a big scope shift just landed mid-delivery — now what?". Given the current engagement state and a change event, it re-reads the risk mix and the intensity dials, tests whether the change crosses the product-vision line (re-opening Direction qualification: proceed / redirect / stop), re-staffs against the bench, and produces a confidence-tiered re-estimate and the next slice to attack — every claim traced to a named element of the model. It reads the model through the capability-model MCP server rather than carrying a copy, so it cannot go stale; if the server is unavailable it stops rather than reconstructing the model from memory. It proposes; owners decide. It never silently re-scopes, re-prices, or proceeds across a product-vision change.
|
judgment-kit
Run the judgment pass before any interface work. Given a brief, a data source, or an existing workflow, refuse to produce labels, navigation, or actions until five things are named — activity, participant, decision, outcome, disclosure boundary — and emit a handoff contract of product-language responsibilities, approved states, and disclosure and action boundaries. Use before designing or generating any UI, before treating a workflow as ready to build, and before installing any generated candidate. Not a design system; it runs upstream of one.
|