Extraction · Document conversion

1,000 handwritten pages into structured Excel

A thousand pages of handwritten PDF records converted into a structured, searchable Excel workbook — delivered before the deadline.

1,000Pages processed
ExcelStructured output
EarlyDelivered

The problem

The client held a thousand pages of handwritten records as scanned PDFs. The information was valuable and completely unusable: not searchable, not sortable, not countable, and impossible to import anywhere.

Handwriting is where automated extraction tends to fail. Optical character recognition handles printed text well and handwriting badly, particularly when:

  • Several different people's handwriting appears across the same set.
  • Scans are skewed, faint, or photographed rather than scanned.
  • Forms were filled inconsistently, with entries crossed out or written in margins.
  • Numbers are ambiguous — a 1 that could be a 7, a 0 that could be a 6.

An OCR pass on material like this produces output that looks complete and is quietly wrong, which is worse than no output at all.

What we did

  1. Define the target schema first

    Agreed the exact columns and formats before any transcription started, so a thousand pages were captured one consistent way rather than reconciled afterwards.

  2. Transcribe manually, with the ambiguity flagged

    Human transcription against a fixed schema. Where a character was genuinely unclear, the cell was flagged rather than guessed, so the client could see exactly which values needed their judgement.

  3. Standardise as we captured

    Dates to one format, names to one convention, numbers as numbers rather than text. Cheap during capture, expensive afterwards.

  4. Quality-control pass against the source

    A sample re-checked page against page, verifying the transcription matched the original rather than merely looking plausible.

The result

One structured Excel workbook covering all 1,000 pages, delivered before the deadline, with uncertain values flagged rather than silently guessed.

The flagging is the part clients tell us they value. A transcription that admits to fifteen uncertain cells is more useful than one that hides them, because you know precisely where to look.

Last reviewed 28 July 2026. Figures are taken from the delivered project record. Client names are withheld where no permission to name has been given.

Questions

About this project

Why not just use OCR software?

OCR is excellent on printed text and unreliable on handwriting, especially with mixed hands or imperfect scans. Its failure mode is the dangerous one: it returns confident output that is wrong. For handwritten source material we transcribe manually and flag anything genuinely ambiguous.

What formats can you work from?

PDFs, scanned documents, photographs, images, websites and handwritten notes. If it can be read by a person, it can be transcribed - the question is only how long it takes.

How do you handle unclear handwriting?

We flag it rather than guess. The delivered file marks every uncertain value so you can check those specific cells against the original, instead of having to trust the whole file equally.

What does document extraction cost?

Quoted by page count and complexity, after seeing a sample. Clean printed forms are far cheaper per page than mixed handwriting on faint scans, so we ask for a representative sample rather than the easiest page.

Have a similar problem?

Send a sample of your data. We will tell you what is wrong with it, what it would take to fix, and what it costs — before you pay anything.

Get a free assessment WhatsApp