A research platform for a small group of non-technical people
Blog post #89
This is the project that has given me the most value so far, and it started as a folder of papers. The material is not public yet, so I leave out names, places and anything that would point to the people involved. The method is the interesting part anyway.
What changed since last log
- The archive stopped being a pile of files I had read and became a place where five or six people work together.
- Everyone has a different job and a different language. A librarian decides how a dialect should sound in Swedish. A journalist and a professor check the German. The family checks the facts. I make sure they can all do that without installing anything.
What shipped
- A bilingual site (German first, Swedish second) with a language switch on every page.
- Review rooms. Each text sits next to its original image or recording. You correct the text directly in the box, press send, and mark it approved, uncertain or reviewed.
- A shared database behind the site (Supabase). Comments, corrections and statuses are saved there, so everyone sees what everyone else did. A shared access code keeps it closed.
- A translation workspace. Instead of translating a whole book in one style, I made one short passage in three styles for the librarian to choose from. The choice becomes the style guide.
- Audio for old texts. Poems read in my cloned voice, labeled as such, and never presented as the author’s voice.
- Illustrations to explain the setting, generated with the image tool in ChatGPT’s Codex CLI, labeled as illustrations and never as photographs.
- A glossary page where each dialect word shows a clean example, the original page image, and a yellow mark on the exact word.
- A project plan page that says what is done, what is in progress, who we are waiting for and what is next.
What’s working
- Two subject experts guided the material. They knew what mattered and what was a detour. That saved more time than any script.
- Mail straight from the prompt. I sent the invitations, the access code and suggested messages to forward from the same conversation where I built the pages. No copy and paste between tools.
- The original next to the text, always. Nobody has to take my word for what a page says.
- Marking what is machine-made. OCR text, machine translation, voice and images are all labeled. People trust the rest more because of it.
- Small decisions first. A three-style sample is cheap. A rewritten 80,000-word book is not.
What’s unclear or broken
- The text box in the review room was invisible in Swedish. A language attribute I put on it matched the rule that hides the other language. I only saw it because someone sent a screenshot.
- Scanned images that look right in my browser can be upside down elsewhere. I now rotate the files themselves instead of relying on styling.
- My first glossary examples came from a bad OCR run and looked like noise. Better text and the page image fixed it, but I should have checked one card before publishing all of them.
- Two poems had no scan because the file had no number in its name, so my catalog skipped it.
- A phone test failed for reasons I can only guess at, and I have no iPhone to check.
Decisions made
- One source of truth for the text, and everything else is a view of it.
- No names and no personal data in public posts until the people involved say yes. The site itself is unlisted and the comment tools need a code.
- A human decides taste. The tool proposes, the librarian chooses.
- Neutral colors and a serif face. A rounded font and a loud accent color looked like a toy, and these are serious papers.
Tooling & process
- Claude Code for the building and the mail, Tesseract and Whisper for reading and transcribing, ElevenLabs for the voice, Codex for the pictures, Supabase for the shared data, and a plain static site on top.
- The biggest lesson: the technical part is the easy half. The value is in giving each person one clear place to do one clear thing, and in making the original visible next to everything they are asked to judge.
— Stefan