Substrate Audits Before AI Features

Substrate audits before AI features

Before building an AI feature, check whether the non-AI substrate can already support the promise. The useful question is not "could an LLM do this?" It is "does the system have the data, lifecycle hook, intent bit, citation shape, privacy gate, and client surface that would make this feature true?"

This came out of the mxr AI email roadmap. The feature ideas were good: pre-send safety, owed replies, archive questions, decision logs, briefings, send-time advice. The first code-truth pass found that some foundations were not ready yet: cadence fields were placeholders, outbound cc and bcc were not part of contact refresh, reply-all intent was not persisted on drafts, and safety needed to run before a draft entered Sending.

By the later code pass, those gaps had mostly been fixed. Contacts now aggregate outbound to, cc, and bcc; reply cadence comes from reply_pairs; DraftIntent persists reply_all; the send path runs safety before the status CAS; semantic hits carry chunk metadata; and the web/docs surface exists beside CLI and TUI. The learning is not that the original plan was wrong. The learning is that the plan only became implementable once the boring substrate became real.

The audit checklist

Run this before green-lighting any AI or "smart" feature:

If any answer is "not yet," that is Track 0 work. Do the substrate first.

The stopping rule

Track 0 is complete when the next user-visible slice is safe to ship. It is not an invitation to finish every piece of operational machinery before returning to the product.

A CourseLit launch made the boundary obvious. The deployment work caught real restore and migration failures. It also grew enough to delay the course and landing page until I stopped the work and asked why the machinery had become the whole project. The correction was to keep the safety fixes and run the visible product work alongside them.

After the next slice is safe, more substrate work needs a concrete failure to justify delaying it.

Why this matters

AI features often hide missing product infrastructure. An LLM can produce a plausible answer even when the system did not preserve the intent bit or evidence row that would make the answer trustworthy. That is the dangerous version of cleverness: it makes the demo work while the product semantics stay weak.

The stronger pattern is slower at the beginning and faster later. Build the substrate, then layer intelligence on top. Once the data is real, the AI layer becomes thin: retrieval, synthesis, validation, display. Without it, every feature becomes a special case.

How this generalizes

In each case, the AI feature is the last mile. The substrate earns its place by making that mile safe. Once it does, further substrate work needs a concrete failure to justify delaying the feature.

See also