Skip to content
Lamka LabsTranscription

Resources

What makes a recording difficult to transcribe

6 min read

Difficulty is about intelligibility, not length

A three-hour conference recording with a good microphone and one speaker at a time can be more straightforward than a forty-minute meeting captured on a laptop across a large table.

Length affects cost, because cost is per audio minute. Difficulty affects effort, turnaround, and sometimes which tier a file belongs in. They are separate things.

What makes audio difficult is anything that forces a reviewer to replay a passage. Every replay is time, and on a badly recorded file the replays never stop.

Microphone placement and room acoustics

Distance is the single biggest factor. The further a microphone sits from a speaker, the more of the room it records relative to the voice. A recorder in the middle of a boardroom table mostly records the table.

Hard surfaces reverberate. Glass-walled rooms, bare floors and largely empty spaces produce an echo that smears consonants together, and consonants carry most of the information needed to tell similar words apart.

Continuous background noise is worse than intermittent noise. Air conditioning, traffic hum, a projector fan or a nearby refrigerator sits underneath the speech for the entire recording. People in the room stop noticing it within minutes. The microphone never does.

Phone and video-call audio is compressed, and compression discards exactly the frequency detail that distinguishes one consonant from another. A call recording is usually harder than an in-person recording made in the same conditions.

Multiple speakers and crosstalk

This is the hardest common problem. When two people speak at the same time, there is often no recoverable answer to what each of them said, no matter who listens or how many times they replay it.

Attribution degrades before intelligibility does. Even where the words are clear, reliably working out which of five voices produced them is difficult, particularly when speakers have similar voices or the recording is mono.

Focus groups are the usual example. A single omnidirectional microphone in the centre of a room full of people who have been asked to talk to each other produces every condition above, by design.

The practical consequence is that crosstalk-heavy recordings carry more inaudible and crosstalk tags. That is the honest outcome. The alternative is a transcript that invents an answer and does not tell you where.

Accents, specialist vocabulary and proper nouns

These get grouped together and should not be. They are three different problems with three different solutions.

An accent is a listening problem. A reviewer familiar with a speech pattern hears it without difficulty. The solution is having listened to a great deal of speech, which is most of what experience amounts to.

Specialist vocabulary is a knowledge problem. A reviewer who works in a field routinely knows the term, knows how it is spelled, and knows when a word that sounds close is wrong. This is what the specialist tier at $1.49 per audio minute exists for.

Proper nouns are a verification problem, and often nobody outside the room can solve it. A surname, a company, a place or a case reference may be perfectly audible and still have several plausible spellings. This is the one where you can help us directly.

What you can do before you record

Move the microphone closer. Closer to the speaker beats better equipment further away, consistently.

Choose a room with soft furnishings if you have the option, and switch off what you can, particularly air conditioning and fans.

Ask participants to avoid speaking over one another and to leave a beat after each other. In a focus group, this is worth saying explicitly at the start rather than assuming.

Ask each speaker to say their name at the beginning. Thirty seconds of introductions makes attribution dramatically more reliable across the whole recording.

Send a list of names and technical terms with the file: participants, organisations, products, case references, anything unusual. This removes the largest single category of avoidable error, and it costs you a few minutes.

How difficulty affects the quote

Difficult audio can move a file to the specialist tier, extend the turnaround beyond the standard two to three days, or both.

We assess this at intake and tell you before work begins, so you can decide whether to proceed, re-record, or accept a longer turnaround.

We do not revise a rate after delivery. If we quoted it, that is the price.