verify
run before a done, fixed, or green claim crosses a gate; each claim gets the command that checks it and the output that could falsify it, or is stamped UNVERIFIED
Goes through every done, fixed and green in a report and pairs it with the command that checks it and the output it produced, or marks it unverified. The table travels with the claim.
structure generate skill verifySkill — verify
Invoke in any agent conversation with /verify before reporting work finished: a claim to
command to output pass over every "done", "fixed", "green", and "ready" in the report. The
deliverable is the verification table, and the table travels with the claim — whoever approves
reads the evidence, never the adjective. The steps are a starting opinion; edit them until they
read the way this company checks its work.
When to use. Before any completion claim crosses a gate: a review waiting on approval, a release about to cut, a report about to land in memory. After a fix, before the word "fixed". Whenever a sentence asserts a state of the world that a command could check.
The Iron Law. Claim → command → output, no skippable steps. Every claim maps to the exact command that checks it and the output that would falsify it. A claim with no command is UNVERIFIED, and the table says so plainly — in the verdict column, not in a footnote. Asserting a verdict without running its command is not a shortcut taken, it is a false sentence written: the report now states something about the world that nobody checked, and every reader downstream spends it as though it were true. There is no step small enough to skip and still call the result evidence.
The deliverable. One row per claim, no prose verdicts:
| claim | command | output | verdict |
|---|---|---|---|
| engine typechecks | deno check packages/engine/src/verify.ts | exit 0, no diagnostics | VERIFIED |
| sensor warns on fake green | deno test -A packages/engine/test/verify_test.ts | ok, 16 passed | VERIFIED |
| TUI unaffected | — | — | UNVERIFIED |
Verdicts are a closed set. VERIFIED: the command ran, the output is shown, and the output supports the claim. FAILED: the command ran and the output contradicts the claim — say so and stop; a FAILED row is a finding, not a drafting problem. UNVERIFIED: no command ran, and the row says why. Exit codes are read unpiped — a pipe returns the last command's code, and the green it manufactures is fake.
Rationalizations, pre-refuted. The excuses arrive on schedule; the answers are already written:
| the excuse | the answer |
|---|---|
| "it passed earlier in the session" | run it again; earlier is not now, and the file has changed since |
| "the change is trivial" | trivial changes ship broken releases; the command costs seconds |
| "the full suite is too slow" | verify the targeted file and write the narrowed scope into the row |
| "I watched it work while developing" | paste the output; output nobody can read is output nobody has |
| "trust me, it works" | see the refusal below |
Process fit. The table is evidence for the gates this company already runs: attach it to the
work a reviewer approves, cite it in the report a release reads, and let the verdict column say
what is actually known at the crossing. Changes the verification uncovers in gated documents land
as proposals (structure propose), never direct writes — a FAILED row earns a fix or a finding,
not a quiet edit to the claim.
The sensor. structure check computes the mechanical subset of this discipline (the verify
sensor, warn tier): a green verdict whose command or output cell is empty is flagged as fake green.
The sensor only ever demotes — it never promotes UNVERIFIED to green, because promotion takes a
human reading real output. A warning is a reason to look, not a build break; the gate reading the
table decides what it accepts.
The refusal. This skill refuses a trust-me green: handed a claim set with no commands and asked to stamp it verified — "trust me, it works", "we are out of time, mark it green" — it declines and returns the same claims stamped UNVERIFIED, each row naming the command that would have earned the verdict. A claim with no command is UNVERIFIED, whoever is asking; the skill never promotes UNVERIFIED rows on assurance alone, because an unverified claim said honestly is recoverable and a fake green is not.
---
type: skill
name: verify
description: run before a done, fixed, or green claim crosses a gate; each claim gets the command that checks it and the output that could falsify it, or is stamped UNVERIFIED
standing: draft
owner: "{{owner}}"
updated: "{{date}}"
sources: "obra/superpowers verification-before-completion (MIT) — see SOURCES.md beside this file"
---
# Skill — verify
_Invoke in any agent conversation with `/verify` before reporting work finished: a claim to
command to output pass over every "done", "fixed", "green", and "ready" in the report. The
deliverable is the verification table, and the table travels with the claim — whoever approves
reads the evidence, never the adjective. The steps are a starting opinion; edit them until they
read the way this company checks its work._
**When to use.** Before any completion claim crosses a gate: a review waiting on approval, a
release about to cut, a report about to land in memory. After a fix, before the word "fixed".
Whenever a sentence asserts a state of the world that a command could check.
**The Iron Law.** Claim → command → output, no skippable steps. Every claim maps to the exact
command that checks it and the output that would falsify it. A claim with no command is UNVERIFIED,
and the table says so plainly — in the verdict column, not in a footnote. Asserting a verdict
without running its command is not a shortcut taken, it is a false sentence written: the report now
states something about the world that nobody checked, and every reader downstream spends it as
though it were true. There is no step small enough to skip and still call the result evidence.
**The deliverable.** One row per claim, no prose verdicts:
| claim | command | output | verdict |
| --- | --- | --- | --- |
| engine typechecks | `deno check packages/engine/src/verify.ts` | exit 0, no diagnostics | VERIFIED |
| sensor warns on fake green | `deno test -A packages/engine/test/verify_test.ts` | ok, 16 passed | VERIFIED |
| TUI unaffected | — | — | UNVERIFIED |
Verdicts are a closed set. VERIFIED: the command ran, the output is shown, and the output supports
the claim. FAILED: the command ran and the output contradicts the claim — say so and stop; a FAILED
row is a finding, not a drafting problem. UNVERIFIED: no command ran, and the row says why. Exit
codes are read unpiped — a pipe returns the last command's code, and the green it manufactures is
fake.
**Rationalizations, pre-refuted.** The excuses arrive on schedule; the answers are already written:
| the excuse | the answer |
| --- | --- |
| "it passed earlier in the session" | run it again; earlier is not now, and the file has changed since |
| "the change is trivial" | trivial changes ship broken releases; the command costs seconds |
| "the full suite is too slow" | verify the targeted file and write the narrowed scope into the row |
| "I watched it work while developing" | paste the output; output nobody can read is output nobody has |
| "trust me, it works" | see the refusal below |
**Process fit.** The table is evidence for the gates this company already runs: attach it to the
work a reviewer approves, cite it in the report a release reads, and let the verdict column say
what is actually known at the crossing. Changes the verification uncovers in gated documents land
as proposals (`overbot propose`), never direct writes — a FAILED row earns a fix or a finding,
not a quiet edit to the claim.
**The sensor.** `overbot check` computes the mechanical subset of this discipline (the verify
sensor, warn tier): a green verdict whose command or output cell is empty is flagged as fake green.
The sensor only ever demotes — it never promotes UNVERIFIED to green, because promotion takes a
human reading real output. A warning is a reason to look, not a build break; the gate reading the
table decides what it accepts.
**The refusal.** This skill refuses a trust-me green: handed a claim set with no commands and asked
to stamp it verified — "trust me, it works", "we are out of time, mark it green" — it declines and
returns the same claims stamped UNVERIFIED, each row naming the command that would have earned the
verdict. A claim with no command is UNVERIFIED, whoever is asking; the skill never promotes
UNVERIFIED rows on assurance alone, because an unverified claim said honestly is recoverable and a
fake green is not.
type: skill
name: verify
about: "Goes through every done, fixed and green in a report and pairs it with the command that checks it and the output it produced, or marks it unverified. The table travels with the claim."
version: "0.1.0"
Also in the folder
Sources — verifyskills/verify/SOURCES.md
Sources — verify
Provenance for the skill beside this note. This file is attribution, not a template: no generator renders it into a company, and the template-walk pins carve it out by name.
- obra/superpowers —
verification-before-completion(MIT License, Jesse Vincent; https://github.com/obra/superpowers). The Iron-Law framing, the run-it-again discipline, and the rationalization-table idiom (pre-refuting the excuses the agent is observed to make) are adapted from there. Ideas and discipline only: the skill body carries no upstream wording. - Corrected 2026-08-22. This note used to record that the upstream's phrase "skip any step =
lying" was kept verbatim in
SKILL.md, and it was. Under the ratified licensing boundary the documents a template generates arrive clean, so the sentence was rewritten in our own words and the lifted phrase no longer ships. NOTICE is append-only and still carries the earlier entry; the correction is appended there rather than edited over. - Adapted 2026-08-23 under the authorized-skills direction: the verification table as the
deliverable, process awareness (gates read the table; gated changes go through proposals), and
the mechanical subset computed by
structure checkas the verify sensor, warn tier. - Attribution is also recorded in
NOTICEat the repository root, which is append-only.