Substrate Audits Before AI Features
Substrate audits before AI features
Before building an AI feature, check whether the non-AI substrate can already support the promise. The useful question is not "could an LLM do this?" It is "does the system have the data, lifecycle hook, intent bit, citation shape, privacy gate, and client surface that would make this feature true?"
This came out of the mxr AI email roadmap. The feature ideas were good: pre-send safety, owed replies, archive questions, decision logs, briefings, send-time advice. The first code-truth pass found that some foundations were not ready yet: cadence fields were placeholders, outbound cc and bcc were not part of contact refresh, reply-all intent was not persisted on drafts, and safety needed to run before a draft entered Sending.
By the later code pass, those gaps had mostly been fixed. Contacts now aggregate outbound to, cc, and bcc; reply cadence comes from reply_pairs; DraftIntent persists reply_all; the send path runs safety before the status CAS; semantic hits carry chunk metadata; and the web/docs surface exists beside CLI and TUI. The learning is not that the original plan was wrong. The learning is that the plan only became implementable once the boring substrate became real.
The audit checklist
Run this before green-lighting any AI or "smart" feature:
- Data truth: Are the fields populated by current code, or are they placeholders?
- Lifecycle placement: Is there a safe place to run the check before side effects?
- Intent capture: Does the model store user intent, or would the feature infer it later from weak signals?
- Evidence shape: Can the output cite source rows, message ids, chunks, or threads?
- Privacy gate: Does the call respect existing local/cloud policy before prompt construction?
- Client parity: Does the daemon expose it so CLI, TUI, web, and agents can all use the same behavior?
- Doc truth: Do docs describe the current implementation, not the plan that used to be true?
If any answer is "not yet," that is Track 0 work. Do the substrate first.
The stopping rule
Track 0 is complete when the next user-visible slice is safe to ship. It is not an invitation to finish every piece of operational machinery before returning to the product.
A CourseLit launch made the boundary obvious. The deployment work caught real restore and migration failures. It also grew enough to delay the course and landing page until I stopped the work and asked why the machinery had become the whole project. The correction was to keep the safety fixes and run the visible product work alongside them.
After the next slice is safe, more substrate work needs a concrete failure to justify delaying it.
Why this matters
AI features often hide missing product infrastructure. An LLM can produce a plausible answer even when the system did not preserve the intent bit or evidence row that would make the answer trustworthy. That is the dangerous version of cleverness: it makes the demo work while the product semantics stay weak.
The stronger pattern is slower at the beginning and faster later. Build the substrate, then layer intelligence on top. Once the data is real, the AI layer becomes thin: retrieval, synthesis, validation, display. Without it, every feature becomes a special case.
How this generalizes
- A CRM recommender needs real relationship events before it can suggest follow-ups.
- A project-management assistant needs status transitions and ownership history, not just task descriptions.
- A code-review bot needs line-range citations and diff context, not just repository search.
- A support copilot needs ticket resolution evidence, not just similar tickets.
- A virtual try-on product needs ownership checks, upload limits, credit reservation, generation idempotency, and status leases before model quality can be the main question.
In each case, the AI feature is the last mile. The substrate earns its place by making that mile safe. Once it does, further substrate work needs a concrete failure to justify delaying the feature.
See also
- Deterministic Before LLM - use exact local evidence first
- Materialize Only When Forced - avoid speculative caches, but document the path
- Citations Required, Validator Enforces - evidence has to be checkable
- Generated Docs as Drift Defense - docs need their own drift controls
- Mxr - the concrete project where this pattern showed up
- Linganisa - the virtual try-on case where generation safety depends on product substrate