A Neighbor's Grandfather, 1,850 Files and a Novel in Dialect

Blog post #85


A neighbor handed me a folder. Inside: material about his grandfather, a Sudeten German journalist and Social Democrat who wrote anti-war books, had them burned by the Nazis in 1933, and fled to Sweden in 1938. About 1,850 files, 9.4 GB, nearly all of it photographs of paper. A professor and a journalist abroad want to research it and write about it, and none of the three can read Swedish. So the job became: make this readable, searchable and available in German.

What changed since last log

  • A new, very different project: not code to ship but an archive to understand.
  • The first real test of “Claude as a research assistant”: look around a folder, tell me what’s in it, then plan the work.

What shipped

  • An overview page with a timeline of 27 events (each tagged with its source and a warning where sources disagree), a material map by collection, a photo gallery, films and audio you can play in place.
  • OCR for the whole archive. 431 of 479 PDFs had no text layer. A local run of Tesseract (German and Swedish) turned 188 unique documents, about 1,350 pages, into plain text files next to each PDF.
  • A reading view for the novel, as ordinary HTML text with the scanned page beside it, search, and a flag button so the family can mark bad pages.
  • A page that explains the OCR work in plain language, with a progress bar that updates while the job runs.
  • A project plan with seven tracks, risks and seven decisions we need from the family.
  • A first dialect glossary (about 760 occurrences of recurring patterns, with suggested Swedish equivalents), and a catalog of 122 scanned poem pages.
  • One poem read aloud with my cloned voice, in German.

What’s working

  • The pipeline is boring in a good way: render page, try four rotations, pick the one with the most common German words, split two-page spreads, run OCR, keep a log.
  • Typed material came out well. 134 documents scored as good, 41 as usable.
  • Seeing a progress bar move on a page the neighbors can open turned out to be the best way to explain “what is Claude doing right now” to non-technical people.
  • Handwritten letters from the 1930s and 40s are readable by Claude directly from the image, even where Tesseract returns noise.

What’s unclear or broken

  • OCR errors are everywhere in the hard cases: blackletter type, photos of books taken sideways, blurry carbon copies. I flag them instead of hiding them.
  • The diary and many documents are handwritten. That needs a different method, and a human proofreader.
  • The sources disagree on basic facts (birth year 1886 or 1889, book published 1929 or 1930). The timeline says so rather than choosing.
  • Rights and privacy: archive documents, police letters and family details. Nothing goes on the open web until the family has said yes.

Decisions made

  • German is the source of truth. Everything the professor and the journalist get must not depend on a Swedish text.
  • The novel is translated directly from German, with a shared dialect glossary and a style guide. No regional Swedish dialect, because that would move the book to the wrong place.
  • Marked as AI: machine-read text, machine translation and the cloned voice are all labeled as such.
  • A mistake worth writing down: I generated the poem with an older voice model out of habit, and was corrected: use the newest one, eleven_v4, released days earlier. I stored it in memory as a rule, because my own knowledge of models lags behind and the right move is to look the model up, not assume.

Tooling & process

  • Claude Code in the desktop app, Tesseract OCR with the German, Swedish and Fraktur language packs, PyMuPDF for rendering pages, 1Password CLI for the API key, ElevenLabs for the voice.
  • Long jobs ran in the background while I built pages in parallel; a small status file fed the live progress bar.
  • Biggest lesson: the cheap part of OCR is the engine. The work is in knowing why a page failed (rotation, spreads, type style) and fixing the cause once for the whole batch.

— Stefan