Assessment computes the facts and the model only writes them up
A language model handed a ledger produces a confident number, so the assessment tool decides what is true and returns an explicit not-enough-data with a reason wherever a metric cannot be computed honestly.
The risk in “have Claude review my books” is specific and it is not hypothetical: a burn rate off by three weeks, “software is up 40%” because the comparison month is half recorded, a $9,500 anomaly that is the rent. Each is worse than silence, because the customer acts on it.
So the split is that the assessment module decides what is true and the model writes prose about it. Anything derived carries whether the data supports it, and a metric that cannot be computed honestly comes back marked insufficient with a reason — never a number with a caveat attached, because caveats are what get dropped in summarisation.
Three decisions do the work. A month counts only if the recording period spans it end to end; partial months are still reported, because the customer wants to see this month, but never averaged and never used as a comparison. Outliers are judged against an account’s own median with a minimum of four postings, which is what separates a 4x grocery bill from the rent. And the run rate and category shares return nothing at all for a multi-currency book, rather than a sum of unlike things.
It was verified through real bean-query rather than a mock, because the column names, the space-padded amounts and the sign of an income posting are all things a mock would have let us get wrong.