Series: AI-First Software Development with SDD and Harness Engineering Part 1: Spec-Driven Development with OpenSpec (you are here) · Part 2: Harness Engineering, Guides, Sensors.
TL;DR — AI coding agents make code cheap. They don't make correct code cheap. AI-first development closes that gap with two disciplines: Spec-Driven Development (SDD), which makes intent a versioned, reviewable artifact, and harness engineering, which makes the agent's output verifiably correct.
This part covers SDD. We explain what it is and why it matters, set up OpenSpec with Claude Code, and build a real feature end to end: discount codes for a checkout API. Part 2 adds the harness that proves the generated code matches the spec.
Who this is for: architects, tech leads and senior engineers deciding how their teams should build software with AI agents. That means more than whether to use them.
Companion repo: shoplite-proj, the full worked example, with specs, code, tests and harness.
1. Why AI-first development needs SDD and harness engineering
The problem: “vibe coding” doesn’t scale
Most teams adopt AI coding agents the same way. Someone opens a chat, describes a feature in a paragraph, and the agent produces a few hundred lines of plausible code. For a prototype this feels magical. For a production system owned by several teams over several years, it breaks down in predictable ways:
The agent "fills in the blanks" with assumptions nobody agreed to → because: Requirements live only in a prompt, and the prompt is ambiguous
Two engineers get two different implementations of the same ask → because: No shared, reviewable artifact exists before code
A reviewer gets a 1,200-line PR and can't tell whether it is correct → because: No statement of intended behavior exists to check the code against
The happy path works but the edge cases were silently skipped → because: Nothing checks which promised behaviors are actually tested
Six months later nobody knows why a rule exists → because: he rationale lived in a chat session that is gone
These symptoms come from two distinct gaps:
An intent gap. The agent doesn’t reliably know what to build, because intent isn’t a first-class, versioned artifact.
A verification gap. Even with clear intent, a probabilistic agent doesn’t reliably build it correctly, and humans can’t review AI-speed output line by line.
SDD closes the first gap. Harness engineering closes the second. Each is weak on its own: a perfect spec with no sensors is a wish list, and a strong test suite with no spec checks the wrong thing very rigorously. Together they form the basis of AI-first software development.
What is Spec-Driven development?
Spec-Driven Development is a way of working in which a human-reviewed, machine-readable specification of observable behavior is written before implementation. That specification is then the input to code generation, the yardstick for verification, and the living documentation of the system afterwards.
Three properties separate SDD from “write a design doc first”:
The spec is structured enough for machines. Requirements and scenarios follow a fixed grammar that tools can parse, validate, diff and merge.
Specs change through deltas. You don’t rewrite the spec. You propose ADDED / MODIFIED / REMOVED requirements and review them like a code diff.
The spec is verifiable. Every requirement carries concrete WHEN/THEN scenarios with exact values, so conformance can be checked. Each scenario maps to a test, and a reviewer can judge a diff against it.
Why now? Because the bottleneck moved
Before AI agents, typing code was expensive enough that a human held the whole intent in their head while writing it. Now generating code costs almost nothing, and the scarce resources are:
Clarity of intent. Does the agent know exactly what to build?
Verification capacity. Can we confirm it built the right thing, at the speed it builds?
That moves where human attention pays off. The later a misunderstanding is found, the more it costs to undo:
SDD moves human effort to the left, where a 60-line spec delta can be reviewed in minutes. Harness engineering automates the right-hand side, so that by the time a human sees code, the mechanical questions are already answered.
2. Setup, process flow and the AI agent
What is OpenSpec?
OpenSpec (by Fission AI) is a lightweight, open-source SDD framework. It has two parts:
A CLI (
openspec) that scaffolds changes, validates spec grammar, reports progress and merges deltas into the source of truth.Agent integrations (slash commands and skills) for 40+ AI coding tools, which teach the agent the workflow.
It is deliberately lightweight: no server, no database, no API keys. Everything is Markdown in your repo, reviewed in your normal PR flow.
Versions used in this guide: OpenSpec 1.14.0, with Claude Code as the agent.
The running example
To keep things concrete, the rest of this post follows one example, which section 3 builds end to end:
ShopLite is a small TypeScript checkout API. It already has one capability described in OpenSpec:
checkout.The change. Marketing wants promo codes, so we create an OpenSpec change named
add-discount-codes. It adds a new capability,discount-codes, and modifiescheckoutso order totals subtract the discount.The timeline. The change was proposed, reviewed, implemented and archived on 2 October 2026, which is why that date appears in folder names below.
The sections that follow show snapshots of this change at different moments, and each one says when it was taken.
The AI agent: Claude Code
We use Claude Code (Anthropic's agentic coding tool, available in the terminal and in VS Code / JetBrains) as the implementation agent. It's a good fit for AI-first development because it exposes every harness extension point we need:
Harness needClaude Code featureWorkflow guidesSlash commands and skills (OpenSpec installs both)Always-on conventionsCLAUDE.md project memoryFast, deterministic sensors in the agent’s loopHooks (scripts that run on agent events and can feed failures back)Independent inferential sensorsSubagents with their own instructions and tools
OpenSpec is agent-agnostic, though. When you install it ( sec: 2.4), openspec init --tools cursor,github-copilot,codex generates the equivalent integration for other tools, and the specs themselves don’t change. Choose your agent for its harness extension points as much as for its model.
Installation and setup
Run these inside the ShopLite repository:
# Prerequisite: Node.js 20.19+
npm install -D @fission-ai/openspec # pin it per repo so CI uses the same version
npx openspec init --tools claude # or run interactively and pick your toolsInstall it as a dev dependency and run it through npm scripts or npx inside the repo. A different, unrelated package called openspec exists on npm, so a bare npx openspec on a machine without the dependency installed fetches the wrong tool.
What init creates:
shoplite-proj/
├── openspec/
│ ├── config.yaml # project context + per-artifact rules
│ ├── specs/ # source of truth (empty on day one)
│ └── changes/
│ └── archive/ # completed changes land here
└── .claude/
├── commands/opsx/ # slash commands: propose, explore, apply, archive, etc
└── skills/ # matching skills Claude Code can auto-invoke
├── openspec-propose/SKILL.md
├── openspec-apply-change/SKILL.md
└── ...The core mental model: specs vs. changes
This is the most important architectural idea in OpenSpec:
openspec/specs/is the source of truth: how the system behaves today, organized by capability.openspec/changes/<change>/is a proposal: how the system should behave after this change, expressed as deltas against the source of truth, plus the reasoning and plan.Archive merges the deltas into
specs/and moves the change folder tochanges/archive/YYYY-MM-DD-<change>/.
Here is add-discount-codes at the moment it is archived:
Architecturally, this is the same pattern as database migrations or Terraform plans. You never edit production state by hand. You propose a diff, review it, apply it, and keep the history. Reviewers see behavioral diffs ("discount no longer applies to shipping") instead of having to reverse-engineer intent from code diffs.
The workflow: explore → propose → apply → archive
You drive OpenSpec through the slash commands that openspec init installed into your agent. Four of them make up the core workflow, and two more help along the way:
There's a human review gate after propose and a PR gate at the end:
Two notes on the diagram:
openspec validate --strict(step 9) checks that the change’s spec files follow the grammar in §2.8. It never looks at code.Steps 13 and 19 (hook sensors and the CI sensor suite) are the harness. Part 1 shows where they sit; Part 2 builds them.
And the lifecycle of a single change:
Three transitions matter more than they look:
Propose stops before code. OpenSpec’s propose command tells the agent explicitly: “This workflow creates planning artifacts only… Do not start implementation in the same response.” That creates a review gate that an eager agent can’t skip.
Implementing → Proposed. When implementation reveals a gap in the spec, you fix the spec first and then the code. This one habit is what stops spec drift.
Archive happens inside the feature PR. The spec update and the code that implements it merge atomically. If you archive after merging,
mainbriefly has code its specs don’t describe, and the archive needs a second PR.
Follow this link for continuation.









