Why Vibe Coding Breaks at Scale: The Case for Harness Engineering
In early 2024, the software engineering industry embraced a seductive new rhythm: "vibe coding." You open a chat window, describe an intent in plain English, watch an LLM generate 400 lines of code across three files, hit reload, and if the UI doesn't crash, you commit.
It feels magical. It feels like 10x leverage. Until month three.
The Entropy Trap: Speed Without Structure
Vibe coding optimizes for instant momentum, not long-term maintainability. When you ask an AI assistant to build an isolated greenfield script, it excels because the entire context fits within a single prompt window. But real production software does not live inside an afternoon chat session. It lives across quarters, hundreds of commits, evolving API contracts, and teams of engineers.
As the codebase expands, four critical failure modes inevitably emerge:
- The Ephemeral Reasoning Void: The crucial architectural decisions — why a specific database abstraction was chosen, what concurrency hazards were mitigated, why an alternative design was rejected — disappear the second the chat tab closes. The code remains, but the reasoning is gone forever.
- "Done" Without Receipts: The assistant announces that the task is finished. But did it actually run the regression suite? Did it check whether changing that shared struct broke a downstream consumer? In vibe coding, "done" is a polite promise until production breaks.
- Context Amnesia & Token Burn: Every new chat session starts completely from scratch. The assistant re-reads the entire repository, burns hundreds of thousands of tokens, introduces subtle stylistic inconsistencies, and slows down.
- Silent Drift Across Repositories: Modern products rarely live in one folder. A change in a contracts library quietly breaks the backend API, which silently breaks the CLI. Without cross-repo verification, team coordination collapses back into frantic Slack messages.
The core problem: AI coding assistants generate code exponentially faster than human engineers can read and verify it. Without an operating harness, AI assistants become firehoses of software entropy.
The Solution: Harness Engineering
The alternative to vibe coding is not going back to manually typing every line of boilerplate. The alternative is Harness Engineering.
Instead of letting an autonomous assistant run unconstrained and hoping for the best, Harness Engineering surrounds the AI with a deterministic, fail-closed operating harness governed by three non-negotiable principles:
1. Hard Human Checkpoints, Not Post-Hoc Reviews
Nothing gets built without plan approval. Nothing ships without evidence verification.
- Gate 1 (Pre-Build): The assistant is strictly prohibited from writing code until a human reviews and approves the architectural plan and test commitments.
- Gate 2 (Pre-Release): The assistant cannot finalize or push code without producing cryptographic, reproducible test receipts.
2. Persistent Project Memory
Your project needs a durable memory stored with the project — not trapped inside a vendor's proprietary web UI:
- A single live snapshot (
handover.yaml) that records active tasks, open decisions, and the single next action. - An append-only journal (
SLICE_LOG.md) that documents why each change was made and what proved it.
A fresh assistant session reads six lines and resumes in seconds with 85% fewer tokens burned.
3. Proof Over Promises
"It compiled on my machine" is not evidence. A true harness executes the agreed verification checks, computes cryptographic hashes over the test outputs, and binds the verification receipt to the Git commit. If evidence is missing, pre-push hooks fail closed.
Bring Harness Engineering to Your Project
Oystro adds deterministic human approval gates and persistent memory to AI coding agents, so your codebase stays reviewable and maintainable for years to come.