build
reach for this before writing implementation code — a new behavior, a bug fix, or a refactor whose outcome a test could catch
Test first, then the code, and every green cites its evidence. Reach for it before writing implementation, most of all when the fix looks too obvious to test.
structure generate skill buildSkill — build
Invoke in any agent conversation with /build, before the implementation exists. One discipline,
two halves, always in order: the failing test comes first, and every green cites its evidence. The
steps are a starting opinion; edit them until they read the way this company builds.
When to use. Before writing implementation code — a new behavior, a bug fix, or a refactor whose outcome a test could catch. Especially the moment the fix looks too obvious to test: that is the moment somebody would otherwise improvise, and the obvious fix that was wrong is the expensive kind.
The Iron Law. No implementation before a failing test, and no green grades itself. A test written after the code passes against whatever the code does, bugs included; a green nobody watched go red proves only that the test can pass.
The loop.
- Read what this company blocks a change over:
departments/engineering/coding-standards.md, its "what a change has to clear before it lands" list. This loop feeds that gate; it never replaces it. - Write the test that fails for the reason the change exists, named after the claim it pins. One claim, one test — a binding claim with no test file named is an opinion.
- Run it and watch it fail. Quote the failing output. A test that passes before the change tests nothing; a test failing for the wrong reason — a typo, a missing import — is not red, it is broken, and fixing it comes before step 4.
- Write the least implementation that turns it green. Resist the extra behavior no test asked for; it is unverified by construction.
- Run again and cite it: the test file, the command, the output, the exit status. A claim that was not run is NOT RUN — written exactly like that, never rounded up.
- Refactor only on green, and re-run after. The suite is what makes a refactor a refactor instead of a rewrite.
- Land through the gate, never around it: a change to a gated path goes out as a proposal
(
structure propose), and the evidence — test names, commands, quoted output — rides in the proposal, so the reviewer approves evidence, not confidence.
The rationalization table. Every excuse below has been observed in the wild. The reality column is the answer, pre-written, so the moment of temptation is never the moment of debate.
| Rationalization | Reality |
|---|---|
| "It compiles, so it works" | Compiling proves syntax. The bar is behavior, and only a test states behavior. |
| "Too simple to break" | Simple code breaks at the same rate; the test costs a minute and holds forever. |
| "I'll write the tests after" | A test written after passes against whatever the code does, bugs included. |
| "I tested it by hand" | A manual run is testimony, not evidence. Nobody can re-run your fingers. |
| "The test is too hard to write" | Hard to test is the design saying where it is wrong. Listen before overriding. |
| "The deadline doesn't allow it" | The deadline allows debugging in production even less. |
| "Just this once" | A discipline that bends when inconvenient was never a discipline. |
| "Just say it works" | Refused — see below. |
Citing green. A claim of green names the test file, the command, and quotes the output, exit status included. "Should pass", "looks right", "the logic is sound" are impressions, not results, and they never stand where a result belongs. The review surface reads evidence, not adjectives; a proposal whose claims carry no commands has earned the reviewer's no before anyone opens the diff.
The refusal. This skill refuses the self-graded green: asked to skip the red step, to write the implementation first "just this once", to report green without running, or to just say it works, it declines and says why. A green that was never red proves nothing, and a cited command is the only form of "it works" this skill is allowed to emit.
---
type: skill
name: build
description: reach for this before writing implementation code — a new behavior, a bug fix, or a refactor whose outcome a test could catch
standing: draft
sources: obra/superpowers (test-driven-development, MIT) · mattpocock/skills (tdd, MIT)
owner: "{{owner}}"
updated: "{{date}}"
---
# Skill — build
_Invoke in any agent conversation with `/build`, before the implementation exists. One discipline,
two halves, always in order: the failing test comes first, and every green cites its evidence. The
steps are a starting opinion; edit them until they read the way this company builds._
**When to use.** Before writing implementation code — a new behavior, a bug fix, or a refactor
whose outcome a test could catch. Especially the moment the fix looks too obvious to test: that is
the moment somebody would otherwise improvise, and the obvious fix that was wrong is the expensive
kind.
**The Iron Law.** No implementation before a failing test, and no green grades itself. A test
written after the code passes against whatever the code does, bugs included; a green nobody watched
go red proves only that the test can pass.
**The loop.**
1. Read what this company blocks a change over: `departments/engineering/coding-standards.md`, its
"what a change has to clear before it lands" list. This loop feeds that gate; it never replaces
it.
2. Write the test that fails for the reason the change exists, named after the claim it pins. One
claim, one test — a binding claim with no test file named is an opinion.
3. Run it and watch it fail. Quote the failing output. A test that passes before the change tests
nothing; a test failing for the wrong reason — a typo, a missing import — is not red, it is
broken, and fixing it comes before step 4.
4. Write the least implementation that turns it green. Resist the extra behavior no test asked for;
it is unverified by construction.
5. Run again and cite it: the test file, the command, the output, the exit status. A claim that was
not run is NOT RUN — written exactly like that, never rounded up.
6. Refactor only on green, and re-run after. The suite is what makes a refactor a refactor instead
of a rewrite.
7. Land through the gate, never around it: a change to a gated path goes out as a proposal
(`overbot propose`), and the evidence — test names, commands, quoted output — rides in the
proposal, so the reviewer approves evidence, not confidence.
**The rationalization table.** Every excuse below has been observed in the wild. The reality column
is the answer, pre-written, so the moment of temptation is never the moment of debate.
| Rationalization | Reality |
| --- | --- |
| "It compiles, so it works" | Compiling proves syntax. The bar is behavior, and only a test states behavior. |
| "Too simple to break" | Simple code breaks at the same rate; the test costs a minute and holds forever. |
| "I'll write the tests after" | A test written after passes against whatever the code does, bugs included. |
| "I tested it by hand" | A manual run is testimony, not evidence. Nobody can re-run your fingers. |
| "The test is too hard to write" | Hard to test is the design saying where it is wrong. Listen before overriding. |
| "The deadline doesn't allow it" | The deadline allows debugging in production even less. |
| "Just this once" | A discipline that bends when inconvenient was never a discipline. |
| "Just say it works" | Refused — see below. |
**Citing green.** A claim of green names the test file, the command, and quotes the output, exit
status included. "Should pass", "looks right", "the logic is sound" are impressions, not results,
and they never stand where a result belongs. The review surface reads evidence, not adjectives; a
proposal whose claims carry no commands has earned the reviewer's no before anyone opens the diff.
**The refusal.** This skill refuses the self-graded green: asked to skip the red step, to write the
implementation first "just this once", to report green without running, or to just say it works, it
declines and says why. A green that was never red proves nothing, and a cited command is the only
form of "it works" this skill is allowed to emit.
type: skill
name: build
about: "Test first, then the code, and every green cites its evidence. Reach for it before writing implementation, most of all when the fix looks too obvious to test."
version: "0.1.0"
Also in the folder
Sources — buildskills/build/SOURCES.md
Sources — build
Provenance for the skill beside this note. It rides the skill's own directory and no generator renders it into a company: the library walk exempts it by name, because attribution to OUR upstreams is this repository's obligation and never a document scaffolded into somebody else's company. The formal attribution is the append-only NOTICE at the repo root.
Adapted from two MIT-licensed skill libraries. Ideas, structure and discipline — no verbatim text:
- obra/superpowers (Jesse Vincent) — the test-driven-development skill: the Iron-Law framing, the watched red (see the failure and its reason before implementing), and the rationalization-table idiom of answering the known excuses inside the skill.
- mattpocock/skills (Matt Pocock) — the tdd skill: red-green-refactor as a minimal, prose-only discipline rather than a toolchain.
Ours, not adapted: the loop reading this company's own coding standards before it starts, the rule that a binding claim names its test file, green cited as claim → command → output with the exit status, and gated output landing as proposals through the review surface.
Where this note came from. The attribution used to be a **Sources.** section inside
SKILL.md, which meant it travelled into every scaffolded company — a document somebody else now
owns, carrying our provenance. Reversed 2026-08-22 under the ratified licensing boundary; the
rendered-output sensor in packages/engine/src/licensing_scan.ts now errors if it comes back.