Write the correctness rule before the code
My first AI product reads Singapore's Parliament and summarises it. These are parliamentary transcripts, and one quote in the wrong mouth would sink the credibility of the whole thing. So the rule came first, before any design existed. Every claim has to show the exact words it came from, in a place anyone can point at.
The failure mode is quiet. A summary that reads well is the one nobody thinks to check. So the rule is not "the model should quote accurately". The rule is verifiable, or the product does not ship.
Three decisions that did the work
Let the model choose, never write. It returns the id of the sentence it means, and the system drops the real text in afterwards. A model can be wrong about which sentence matters. It cannot invent one, because it never types any.
Use the boring mechanism where it can be checked. A foreign key beat a similarity score at saying which speaker a sentence came from. A plain text match beat a vector at knowing that procedural lines are not substance. Ordinary code that gives the same answer twice beat a model that gives a slightly different one each morning.
Refuse to publish, rather than publish with a caveat. If a check fails, nothing goes out. The failure I care about is not a missing page. It is a wrong one going out while I am not looking.
Two things I would not do
I would not put a model in the citation path. If a published record is assembled the moment you read it, then a citation points at a query. The link still resolves, and what it shows can change underneath you.
I would not ask a model to review its own output. A checker that only sees what the model produced will wave through a near-miss paraphrase, because there is nothing independent to compare it against. Asking nicely for accuracy is not a guarantee, and adding a check afterwards does not create one.
What this means for a business
Most of the value was not in the model. It was in the data model, the tests, and a willingness to drop a row rather than guess at it. That is the work that takes longer and gets skipped, because it does not demo well.
A system like this is mostly a data problem with a model in it, not a model with some data attached.
So the question worth asking of any AI feature is narrow, and it is not about capability. What does it do when it cannot be sure? If the answer is that it answers anyway, the feature is not ready for a decision someone has to defend.
This is the kind of work I take on: systems where being close is not good enough, in a setting where the rules are real. If the correctness bar is the hard part of what you are building, get in touch. The full build, including the two designs that failed, is in the Parsnips case study.