How to Transcribe a Family Interview

I have the recording. How do I turn it into something readable?

An hour of family audio sitting on a hard drive is doing almost nothing. Nobody browses audio. Nobody searches it. Nobody reads a paragraph of it aloud at a birthday because they remembered it was in there somewhere.

Text is what makes a recording usable. It can be searched, quoted, printed, corrected, translated and read by a grandchild in ten minutes who would never sit through the tape.

Transcribing is also the least enjoyable part of this whole undertaking, which is why so many family recordings never become anything. Here is how to make it survivable.

Automate the first pass, always

Do not type an interview from scratch. Speech recognition is now good enough that correcting a machine transcript is several times faster than typing one, even when the machine is wrong often.

What it will get right: ordinary conversational English, at a normal pace, in a quiet room.

What it will get wrong, reliably:

  • Names. Every person, street, town, ship, regiment and shop. Expect all of them to be wrong.
  • Dialect and regional vocabulary, which is often exactly the texture you wanted to keep.
  • Older voices, which tend to be quieter and less crisply articulated. This is the biggest single factor.
  • Overlapping speech, where you spoke while they were still going.
  • Numbers and dates, particularly spoken years.

Which tool matters less than the quality of the audio. A clean recording through a mediocre transcriber beats a poor recording through the best one, which is the whole argument for spending ten minutes on the room and the placement.

One caution worth taking seriously: an automatic transcript means uploading your relative’s voice and their private family history to a company. Read what the service says it does with the audio, and prefer one that processes on your own machine if the material is sensitive.

The correction pass

Work with the audio playing and the text open. Do not read the transcript alone and try to spot errors, because a wrong word that makes sense in context is invisible on the page and obvious the moment you hear it.

Go in this order:

  1. Names first, in one pass. Search and replace once you have each spelling right. This is most of the errors and the fastest to fix.
  2. The parts you cannot make out. Mark them [unclear 14:22] with the timestamp rather than guessing. A guess becomes a fact the moment somebody reads it, and family history is full of confident errors that started as a transcription guess.
  3. Everything else, at normal listening speed.

Expect roughly three to four hours of correcting per hour of audio the first time, and less as you get faster. Do it in twenty-minute blocks. It is dull work and quality drops sharply after half an hour.

How much of the speech to keep

This is the real decision, and it depends entirely on what the transcript is for.

A faithful transcript keeps everything: false starts, repetitions, “you know”, the moment they lost the thread and found it again. This is what an archive wants. It is the truest record of how somebody actually spoke, and it is genuinely hard to read.

A readable transcript removes filler and repetition but changes no words. Cut “um”, cut a repeated half-sentence, keep every word they meant to say. Their grammar stays theirs. Their phrasing stays theirs.

A written-up account is not a transcript at all. It is you, retelling what they said in continuous prose. It reads best and it is the furthest from their voice.

For most families the readable transcript is right, with one rule: when you are not sure whether something is filler or character, keep it. The stray “oh, that reminds me” is often the most alive thing on the page, and you cannot recover it once cut.

Whatever you choose, say so at the top of the document in one line. Somebody in forty years will want to know whether these are exactly her words.

Mark the speakers and the time

Label who is speaking, every time it changes. It looks unnecessary while you are doing it and it is essential the moment a third person joins the conversation.

Put a timestamp at the start of each new topic, not each paragraph. [00:14:30] The shop on Mill Street lets somebody find the passage in the audio and hear her actually say it, which is the thing a transcript can never replace.

Fix the facts separately, and visibly

People misremember. Dates slip, two events merge, a name attaches to the wrong person. You will find things in a transcript that you know are not right.

Do not correct their words. Add a note.

“We moved to the new house in 1961.” [The deeds give 1963.]

Square brackets, clearly yours. The transcript records what they said, which is the historical fact you actually captured. Your correction records what the documents say. Both are worth keeping, and quietly editing the first into the second destroys the only copy of the first.

What to do with it once it exists

Print one copy. This is not sentimental: a printed page is the only version guaranteed to be readable without a working account, a supported file format and a device.

Send a copy to the person who spoke, if they are able to read it. It frequently produces a second round of stories, because reading their own account reminds them of what they left out. It also gives them the chance to strike something out, which they have every right to do.

Keep the audio. The transcript does not replace it. A voice carries timing, accent and hesitation that no page holds, and the file costs nothing to store.

Then think about the form the whole thing eventually takes, which is a separate question worth answering deliberately: what to do with a finished family journal.

If an hour of correcting is more than you have

Transcribe the best twenty minutes. Not the first twenty, the best.

A short, accurate, readable passage that people actually pass around is worth more than four unedited hours nobody opens. You can always come back for the rest, and the audio is not going anywhere as long as it is backed up in two places.