What happens between the file and the answer
Short version for the person paying, then the engineering underneath.
In one sentence
Your staff stop reading a 200-page manual to answer one question, and every answer they get carries the page it came from, so it can be checked in seconds instead of trusted blindly.
A language model on its own is a colleague with a very good memory who has never read your documents. This system is the same colleague sitting with your file open, only allowed to answer from the page in front of them, and required to point at it.
Measured difference on a fixed question set: 15% correct without the document, 100% with it.
The pipeline
Timings are measured on this server right now, not written by hand.
…
…
…
< 30 ms
…
Two decisions worth explaining