Live

Vision-RAG Mining Intelligence System

Mining reports that only exist as scans, made queryable without an OCR pass mangling every table on the way in.

Gemini 2.5 FlashFAISSBM25LangChainFastAPIReact

Data Trapped In Paper

Mining companies produce enormous quantities of reports that exist only as PDF scans: geological surveys, environmental assessments, reserve estimates. The data inside them is highly structured. The format it arrives in is not.

The cost of a slow lookup

Analysts querying across a 300-plus document corpus faced hours of manual search per question. Mining investment decisions depend on rapid synthesis of historical report data, and a single misread reserve estimate can mean a nine-figure error.

Where the difficulty actually sits

Not in retrieval. In extraction. These documents have multi-column layouts, geological tables with merged cells, hand-annotated diagrams, and non-standard unit notation. If the ingestion step destroys the table, nothing downstream can recover it.

A second, quieter problem

Jargon-dense language breaks generic embedding models. 'Proven reserves' and 'probable reserves' are near-identical in vector space and legally distinct in practice, which makes pure semantic similarity actively dangerous here.

Why OCR Loses

The standard pipeline

Run Tesseract over the PDFs, chunk the extracted text, embed with a general-purpose model, retrieve by cosine similarity. This is the default answer and it is the wrong one for scanned tabular documents.

Where it collapses

Tesseract on these 300 DPI scans lands around 60 to 70% character accuracy on mining tables. Headers get misread, numbers in adjacent columns merge, and row structure collapses entirely. The resulting text is too corrupted to embed meaningfully, so retrieval quality is capped by a failure that happened long before the query.

Ingestion To Answer

The pipeline treats parsing as the expensive, high-value step and retrieval as the cheap one, which is the reverse of how most RAG systems allocate effort.

What goes in

300 DPI scansGeological surveysReserve estimatesEnv. assessments

What comes out

Reconstructed tablesCited answersIncremental index

Retrieval Design

  1. 01

    Gemini 2.5 Flash vision instead of traditional OCR

    Gemini reads table structure semantically. It can reconstruct a merged-cell table by understanding the document's visual layout rather than extracting a character stream that destroys spatial relationships. The Flash tier keeps a 300-plus document ingestion pipeline affordable.

  2. 02

    FAISS and BM25 fused with Reciprocal Rank Fusion

    Mining terminology is exact-match sensitive: formation names, chemical symbols, regulatory codes. Dense embeddings blur exactly these. BM25 catches them, RRF merges the two ranked lists, and neither retrieval mode has to be right on its own.

  3. 03

    Append-only index with periodic consolidation

    New reports arrive weekly. Re-indexing the full corpus on every addition would create a downtime window each time. An append-only FAISS index with scheduled consolidation keeps the system queryable while documents are being added.

Running a vision model over every page costs materially more per document than Tesseract, and ingestion is slower. That trade only makes sense because the corpus is read far more often than it is written, and because a corrupted table is not a degraded answer, it is a wrong one.

Stack

Gemini 2.5 Flash

Multimodal vision reconstructs table structure from scanned PDFs that traditional OCR cannot parse. The Flash tier balances quality against the cost of a 300-plus document ingestion run.

FAISS

Sub-millisecond retrieval with no standalone vector database to operate. The corpus fits in memory on the deployment server, so the operational cost of a separate service was not worth paying.

BM25

Exact-match retrieval for terminology that dense embeddings blur together. Complements FAISS through Reciprocal Rank Fusion rather than replacing it.

LangChain

Holds the retrieval chain together, including cross-encoder reranking and citation grounding, so every answer can point back to the page it came from.

FastAPI

Serves the query API and streams long answers back to the frontend, with async handlers so a slow retrieval does not block concurrent analysts.

React

The analyst-facing interface, where a cited answer needs to sit next to the source table it was drawn from.

Where It Landed

Results

95%+ accuracy on domain-specific queries across a 300+ document corpus, with sub-3s retrieval latency and zero-downtime incremental indexing on weekly report additions.

The number I trust least

Accuracy here is measured against analyst-written questions, which skew toward the queries people already knew to ask. The honest gap is performance on questions nobody thought to pose, and that is not something the current evaluation set can tell me.

Open Questions

  • —

    Whether vision parsing should run once at ingestion or be re-run when the model improves. The current design bakes in whatever the model understood on ingestion day, and a better model does not retroactively fix an old chunk.

  • —

    How to evaluate table reconstruction directly, rather than inferring it from downstream answer quality. A wrong number in a correctly-shaped table is invisible to the current metrics.

  • —

    Whether reranking earns its latency on this corpus. It was added on the assumption it would help, and the ablation to prove that never got run.

Solving a similar problem?

I'm open to conversations about production AI systems -agentic workflows, RAG pipelines, or messy integration problems like this one.