2026 · Solo
MIME (Make It More Efficient)
A GenAI system that turns source documents into standardized, source-cited memos, where every financial figure is computed in code and an assertion tracer confirmed zero of them were untraceable to the source ledger.
The idea
Give MIME one finished example report and a pile of source documents. It learns the format from that single example, extracts the facts it actually needs, and writes the same report for the new inputs. No hand-coded template, no starting from scratch every time.
The first target is credit memos, a real problem I found out about at a recent company I worked at: relatively standardized on paper, genuinely tedious in practice.
The metric that actually matters
It’s tempting to lead with a cost number. I’m not going to, because it’s not the strongest thing here.
Code computes every financial figure in the memo. The model only ever writes prose. That’s a structural choice, not a prompt-engineering trick, so numeric error is bounded by extraction error rather than by whatever an LLM feels like asserting that day. I built an assertion tracer that reads every number in a generated memo’s prose and checks it against the source ledger. On the live model, across an 11-case frozen eval: 77 of 77 numbers traced cleanly to the source ledger, at 98.6% field extraction accuracy.
Anyone can make an LLM pipeline cheaper. Being able to say “the model structurally cannot assert a number, and here’s the instrument that verifies it” is a different, harder claim, and in a credit-memo context, at a bank, it’s the entire product question. This is also the one claim on this page that’s fully measured rather than modeled, which is exactly why it leads.
The scaling law, not the multiple
A naive approach re-sends every source document for every section of the report, so cost grows as O(sections × corpus). MIME reads each document once into a compact fact ledger, so cost grows as O(corpus + sections × ledger) instead. That’s the real claim: not a single multiple, but a curve that stays flat while the naive approach climbs.
Important caveat: this curve is modeled, not measured. It’s a cost projection at an assumed 500,000-token corpus, not a benchmark run against real API bills, and I want to be upfront about that rather than let a clean number imply more than it’s earned. On that model, at a 50-document, 15-section memo: about 14× fewer input tokens, and about 11× lower cost than a prompt-cached full-context baseline, the competent baseline, not an uncached strawman. (The uncached comparison is closer to 63×, and I’m deliberately not leading with that number: beating a baseline nobody would actually ship isn’t the interesting result.) Calibrating that projection against a real measured run, with the actual API tokenizer instead of a character-count estimate, is the next thing I want to do here.
Where my own system loses, on purpose
Here’s the finding I like most, and it’s not a win. The naive baseline scored 100% on computation accuracy in the eval. MIME scored 94.4%, lower. The reason: the baseline computed a debt-service-coverage ratio that looked right from two inputs that were actually wrong. A top-line accuracy check would have waved it straight through. MIME’s fact-level checks caught the same case and didn’t produce a confidently-wrong number.
Publishing a metric your own system scores lower on, and being able to explain exactly why that metric is the wrong one to optimize, is worth more than the accuracy number itself.
Built for real providers, not just demos
A capability-aware provider abstraction means Claude can be swapped for a self-hosted model without touching the pipeline. That abstraction earned its keep during testing: an early schema used additionalProperties: true, which the real Claude API rejected outright. The offline mock provider, never having to satisfy a real API contract, didn’t catch it at all. Test doubles have structural blind spots. This is the version of that lesson that costs you a debugging session instead of a production incident.
What it doesn’t do (yet)
Gap detection is 100% reliable at the document level: if a required document is missing, MIME flags it before doing any model work. What it can miss is a document that’s present but empty of the field you needed. The document exists, so nothing gets flagged upfront, and the missing value only surfaces later as a field-level warning instead of an initial gap report. I’d rather report that limit voluntarily than have someone find it in the code first.
What it demonstrates
Turning a document-heavy business grind into a repeatable, auditable pipeline, one where the trustworthy part is measured and the efficient part is honestly labeled as a projection still waiting on a real calibration run.