A Neighbor's Grandfather, 1,850 Files and a Novel in Dialect
Blog post #85
A neighbor handed me a folder. Inside: material about his grandfather, a Sudeten German journalist and Social Democrat who wrote anti-war books, had them burned by the Nazis in 1933, and fled to Sweden in 1938. About 1,850 files, 9.4 GB, nearly all of it photographs of paper. A professor and a journalist abroad want to research it and write about it, and none of the three can read Swedish. So the job became: make this readable, searchable and available in German.
What changed since last log
- A new, very different project: not code to ship but an archive to understand.
- The first real test of “Claude as a research assistant”: look around a folder, tell me what’s in it, then plan the work.
What shipped
- An overview page with a timeline of 27 events (each tagged with its source and a warning where sources disagree), a material map by collection, a photo gallery, films and audio you can play in place.
- OCR for the whole archive. 431 of 479 PDFs had no text layer. A local run of Tesseract (German and Swedish) turned 188 unique documents, about 1,350 pages, into plain text files next to each PDF.
- A reading view for the novel, as ordinary HTML text with the scanned page beside it, search, and a flag button so the family can mark bad pages.
- A page that explains the OCR work in plain language, with a progress bar that updates while the job runs.
- A project plan with seven tracks, risks and seven decisions we need from the family.
- A first dialect glossary (about 760 occurrences of recurring patterns, with suggested Swedish equivalents), and a catalog of 122 scanned poem pages.
- One poem read aloud with my cloned voice, in German.
What’s working
- The pipeline is boring in a good way: render page, try four rotations, pick the one with the most common German words, split two-page spreads, run OCR, keep a log.
- Typed material came out well. 134 documents scored as good, 41 as usable.
- Seeing a progress bar move on a page the neighbors can open turned out to be the best way to explain “what is Claude doing right now” to non-technical people.
- Handwritten letters from the 1930s and 40s are readable by Claude directly from the image, even where Tesseract returns noise.
What’s unclear or broken
- OCR errors are everywhere in the hard cases: blackletter type, photos of books taken sideways, blurry carbon copies. I flag them instead of hiding them.
- The diary and many documents are handwritten. That needs a different method, and a human proofreader.
- The sources disagree on basic facts (birth year 1886 or 1889, book published 1929 or 1930). The timeline says so rather than choosing.
- Rights and privacy: archive documents, police letters and family details. Nothing goes on the open web until the family has said yes.
Decisions made
- German is the source of truth. Everything the professor and the journalist get must not depend on a Swedish text.
- The novel is translated directly from German, with a shared dialect glossary and a style guide. No regional Swedish dialect, because that would move the book to the wrong place.
- Marked as AI: machine-read text, machine translation and the cloned voice are all labeled as such.
- A mistake worth writing down: I generated the poem with an older voice model out of habit, and was corrected: use the newest one,
eleven_v4, released days earlier. I stored it in memory as a rule, because my own knowledge of models lags behind and the right move is to look the model up, not assume.
Tooling & process
- Claude Code in the desktop app, Tesseract OCR with the German, Swedish and Fraktur language packs, PyMuPDF for rendering pages, 1Password CLI for the API key, ElevenLabs for the voice.
- Long jobs ran in the background while I built pages in parallel; a small status file fed the live progress bar.
- Biggest lesson: the cheap part of OCR is the engine. The work is in knowing why a page failed (rotation, spreads, type style) and fixing the cause once for the whole batch.
— Stefan