Why a complete PRD still leaves an execution gap
A PRD can specify product intent, but an AI coding agent still has to infer architecture, repository conventions, migration rules, validation commands and what “done” means. Those inferences are exactly where implementation can drift from the approved technical direction.
Modern spec-driven workflows increasingly separate specification, planning, tasks and implementation. GitHub Spec Kit, for example, explicitly structures agent work as Spec → Plan → Tasks → Implement and adds quality/checklist stages around that flow. OpenAI Codex also recognizes repository-scoped instructions through AGENTS.md. The useful lesson is not to copy a specific tool; it is to make engineering intent durable and inspectable instead of relying on one giant prompt.
The six artifacts an agent needs beyond the PRD
| Artifact | Purpose |
|---|---|
| Architecture decision records | Prevent the agent from silently re-deciding high-level design. |
| Repository map | Tell the agent where truth lives and what instructions apply. |
| Implementation plan | Order work into bounded, reviewable steps. |
| Scoped skills/checklists | Apply domain-specific rules only where relevant. |
| Quality gates | Turn architecture/security rules into completion conditions. |
| Acceptance/golden tests | Verify that generated code converges on intended behavior. |
Why “one enormous AGENTS.md” is also a bad answer
Instructions should be layered. Keep the top-level map concise; put detailed knowledge in focused files that can be versioned and tested. That reduces stale context and makes it obvious which rule governs which task.
Human approval still matters
Destructive database changes, privilege changes, production deployment, security-control changes and final acceptance should have explicit human checkpoints. “Agent-ready” should mean constrained and auditable, not uncontrolled autonomy.
See the linked primary reference and runnable example below.
Separate instructions from verification
A file saying “do not import private modules” records intent. A source check reporting the exact violating file and line supplies narrower, inspectable evidence. Keep both, and explain what the check cannot see. A generated checklist is not an executed test.
Decision → rule → source check → finding → correction
ARCH-001 → public module interfaces → Python AST → private import → API importOur worked example includes a deliberately failing source tree, a corrected tree and controls for unsupported or empty scope. The checker never imports the application. It is not a security sandbox or proof that the agent implemented the product correctly.
A handoff contract worth maintaining
- Name the approved baseline and the owner who can change it.
- Keep the task small enough to inspect its completion evidence.
- Record the validation command and expected failure conditions.
- Escalate unknown evidence rather than converting it into a green status.
- Review and rerun controls when the architecture decision changes.
Primary reference
GitHub Spec Kit documentation describes its Spec → Plan → Tasks → Implement flow. SYSLUME's source check is a separate demonstration, not a claim that using that tool guarantees architecture compliance.