Asking 4,200 pages of manuals questions the technicians wrote themselves
The 120 test questions came from the maintenance technicians, not from questions we picked because we could answer them. The documents include both digital files and skewed scans.
Starting situation
4,200 pages of machine manuals in mixed Thai and English. Technicians had to hunt for the answer themselves every time a machine had a problem — slow, and often unsuccessful.
Constraints on site
The documents include both digital files and skewed scans (which have to be OCR'd before they are searchable), and the 120 test questions came from real technicians, not ones we picked to be easy on ourselves.
What we actually did
PaddleOCR turns the skewed scans into text, which is merged with the digital files, then llama.cpp on an RTX 4090 does the retrieval and answering. Everything runs on the machine — no document ever leaves the network.
Results, in numbers
See the figures above — the correct-answer rate and answer time are still awaiting a repeat measurement before we state real numbers. The 0 KB leaving the network is an architectural fact, not something that needs measuring.
RTX 4090 · llama.cpp · PaddleOCR
Want to talk about a problem like this in your plant?
A free 30-minute call. We will tell you straight if your problem is not a good fit for AI yet.
Talk to us for 30 minutes firstSee all work