Services

Custom OCR engineering

Custom OCR development for documents that off-the-shelf OCR gets wrong

Custom OCR pipeline development for scanned, photographed, and low-quality documents — layout-aware extraction, confidence scoring, vision-model escalation, and ground-truth accuracy benchmarks.

From $1,400

Document parsing & extraction

Document parsing and data extraction, built as production infrastructure

Turn unstructured PDFs, scans, and emails into validated structured data. Custom document parsing pipelines with schema design, confidence scoring, human review queues, and measurable accuracy.

From $1,400

Self-hosted VLM deployment

Deploy vision-language models on your own hardware

Set up open-weight vision-language and OCR models on your own GPUs or private cloud — Qwen3-VL, InternVL, and similar. Sizing, quantization, vLLM serving, throughput tuning, and accuracy validation.

From $1,800

RAG systems

RAG systems that retrieve the right thing — and prove it

Design, build, and evaluate retrieval-augmented generation systems: chunking and indexing strategy, hybrid and reranked retrieval, grounded answer generation, and retrieval evaluation harnesses.

From $1,400

Advisory

Document AI consulting: architecture, build-versus-buy, and accuracy strategy

Independent consulting on document AI architecture, model and vendor selection, accuracy measurement, cost modelling, and build-versus-buy decisions — from an engineer who ships extraction systems.

From $180 / hour

On-premise & private cloud

On-premise and private-cloud document AI deployment

Deploy the full document extraction stack inside your own infrastructure — VPC, private cloud, or air-gapped — with no document leaving your network. Sizing, deployment, and compliance documentation.

From $6,500

Commercial arrangements

If you would rather use the engine than build one, these are the ways to do it.

Services

Document AI, OCR, and RAG engineering services

I build production document AI: OCR pipelines for documents that standard tools get wrong, extraction infrastructure with validation and human review, vision-language models running on your own hardware, and retrieval systems that are measured rather than hoped for. Parse My Statement is the reference implementation, so you can inspect the engineering before you hire me for it.

  • Accuracy measured against hand-written ground truth, never asserted
  • Delivered into your repository and your infrastructure
  • Independent engineer — no vendor commissions, no delivery layer
  • Fixed-scope phases, starting with an audit you can stop after

The common thread: measurement

Almost every document AI project I am called into has the same missing piece. There is a pipeline, there are models, there is a dashboard — and there is no ground truth. Nobody can say whether last month's change helped, because nothing is scored against known-correct values.

That single gap explains most of what goes wrong downstream. Teams tune prompts when the problem is retrieval. They swap OCR vendors when the problem is normalization. They ship a model upgrade that quietly degrades accuracy for three weeks before anyone notices, because the only signal is user complaints.

So every engagement I take starts the same way: build the ground truth, score the baseline, then make decisions against numbers. It is unglamorous and it is the reason these projects finish.

How engagements are structured

Fixed-scope phases. You can stop after any of them.

  1. 1. Audit or feasibility study

    One to three weeks, fixed price. A ground-truth set, a scored baseline, a cost and latency model, and a written recommendation. Some clients stop here because the report answers the question.

  2. 2. Build

    Four to twelve weeks. Iterative development against the benchmark, delivered into your repository from the first week, with accuracy visible at every step.

  3. 3. Hardening and handover

    Retries, timeouts, cost caps, observability, runbook, and training — plus the evaluation harness, so your team can keep scoring changes after I leave.

  4. 4. Retainer, if you want one

    Ongoing accuracy work as document mixes and models change. Entirely optional; the handover is designed so you do not need it.

Typical reasons teams get in touch

  • Extraction accuracy has plateaued below what the business needs
  • A document AI pilot works in the demo and not in production
  • Compliance now requires that documents never leave the network
  • The LLM bill is growing faster than the volume that justifies it
  • A RAG system returns plausible answers from the wrong documents
  • An in-house build has stalled and nobody can say why
  • Statement parsing is needed in the product but not worth a team

Frequently asked questions

How do engagements usually start?

With a fixed-price audit or feasibility study. It produces a ground-truth benchmark, a scored baseline, and a written recommendation you can execute with or without me. That structure exists so you can evaluate the work before committing to a build.

Do you work as part of our engineering team?

That is the preferred mode. I build in your repository, follow your conventions, and submit pull requests your engineers review — because the goal is that your team can extend the system after handover.

What does this cost?

Audits start around $1,400, builds around $3,500, and retainers around $1,800–$2,200 per month. Advisory time is $180 an hour. Each service page carries indicative ranges; final scope and price are fixed in writing after a discovery call.

Can you work under our security and compliance requirements?

Yes. NDA and DPA as standard, work inside your environment where required, and fully on-premise or air-gapped deployment where documents may not reach any third party.

Start with the documents

Send a sample of what you need extracted — ideally the cases that currently fail. You will get an honest read on what is achievable and what it would take.

Or email [email protected]

Related

This page is also available as markdown for AI agents: /services.md · index at /llms.txt. Canonical URL: https://parsemystatement.com/services.