Short answer: Record close to both speakers and ask for consent on the recording itself. Then make a first draft: an AI pass takes minutes, careful typing takes four to six hours per hour of audio. Correct that draft against the audio, add speaker labels and timestamps, and mark unclear passages as [inaudible] rather than guessing.
An hour of ordinary conversation runs to roughly eight thousand words. As audio, not one of them can be searched, scanned or quoted with confidence. As text, all of them can. That single conversion is why transcription sits at the centre of journalism, qualitative research, oral history, podcasting and hiring.
Interview transcription is also the least forgiving kind. A meeting summary can smooth over a garbled sentence. A published quotation cannot. The work that follows is not heroic typing. It is a sequence of small decisions taken in the right order, and the first of them happens before the recorder runs.
Set up a clean recording
Most transcription problems are recording problems that travelled downstream. Distance, echo and background noise raise the error rate of every method, human or machine. The cheapest accuracy you will ever buy is bought in the ten minutes before an interview starts.
The quality of a transcript is mostly decided before anyone types a word — at the moment of recording.
Microphone placement beats microphone price
Put the microphone or phone within arm's reach of both speakers, not at the far end of the table. Halving the distance to a speaker raises their level by about 6 dB against the room. A modest lavalier clipped to a collar will beat an expensive recorder two metres away. Aim the microphone at the mouth, not at the ceiling.
Choose the room, then test it
- Prefer soft rooms. Curtains, carpet and bookshelves absorb reflections. Bare walls and glass create the echo that smears consonants together.
- Make a 30-second test. Record, then listen back on headphones before you begin. It is the only reliable way to catch a rustling jacket, an air-conditioning hum or a failing battery.
- Run a backup. A phone recording alongside your main device costs nothing, and it has rescued more interviews than any other habit.
- Record remote calls locally. Where software allows it, capture each side as a separate track and ask your guest to wear a headset. Two clean local tracks transcribe far better than one compressed call recording.
- Keep the master file. Record to WAV or a high-bitrate MP3, then archive the file exactly as it came off the device. Every later stage works from that copy.
Consent, context and the spelling list
Ask permission to record, and ask for it on the recording itself. Consent and audio then live in the same file. Legal requirements vary by jurisdiction, and some places require the agreement of every person recorded. The professional standard is simpler and stricter: never record covertly.
Oral historians have thought about this longer than anyone. The Oral History Association's best practices treat consent as a process rather than a sentence. In research settings that process has three parts. It needs an information sheet, a clear statement of how long recordings will be kept, and a written record of what the participant agreed to.
While the recorder runs, keep a notepad for what the audio cannot carry. Write down the spelling of every name, company, place and technical term as it comes up. Five minutes of noting spellings saves an hour of detective work later. That list also becomes your custom vocabulary if a machine writes the first draft.
If the recording will pass through cloud software, check retention and training terms before you upload anything sensitive. Interview material often contains personal data belonging to someone who is not you. We set out the questions worth asking in privacy and security in AI transcription.
Choose a transcript style before you type
Decide what kind of transcript you are making first, not halfway through the file. The choice trades fidelity against readability, and the right answer depends entirely on what the transcript is for.
| Style | What is kept | Best suited to | Relative effort |
|---|---|---|---|
| Full verbatim | Every word, filler, false start and repetition | Discourse analysis, legal records, linguistics | Highest |
| Intelligent verbatim | Every substantive word; fillers and stumbles tidied | Journalism, podcasts, qualitative research, business | Moderate |
| Summary notes | Key points plus selected exact quotations | Background interviews, internal briefings | Lowest |
No option is neutral. As the study of transcription in linguistics makes explicit, every transcript interprets. It decides where a sentence ends, which repetition matters, and how an accent is rendered on the page. Careful work is not defined by the style you pick. It is defined by consistency, plus a one-line note recording the convention you used.
Produce the first draft
There are two honest routes to a draft. They differ sharply in cost and speed, then meet at exactly the same place: a correction pass against the audio.
The AI-assisted route
A modern speech engine drafts a clear two-person interview in minutes. Timestamps and speaker turns arrive already in place. Feed it the spelling list from the previous step if the tool accepts a custom vocabulary. A tool such as TRANSCRIPT.YOU handles the mechanical pass so your attention goes to the words that matter. It helps to know how speech-to-text turns sound waves into sentences, because that explains which errors better audio can fix and which it cannot.
The manual route
Slow playback to 70 or 80 per cent and work in loops of a few seconds. Type in passes: words first, punctuation and formatting second. Keyboard shortcuts for pause and rewind matter far more than raw typing speed, and a foot pedal frees both hands. Budget four to six hours per hour of audio, and more for poor recordings or several voices.
Typical working time per hour of clear, two-person audio. Those ranges assume a good recording. A poor one erases the advantage quickly, because repairing a bad draft can take longer than typing a clean one from scratch. The wider economics are set out in our comparison of when a professional transcriber still earns the fee.
The correction pass
The first draft is never the deliverable. The correction pass — audio in your ears, text under your eyes — is where a transcript earns its trust. Work through the whole recording once, stopping wherever a sentence reads too smoothly to be true.
Machine drafts fail in a specific way. They rarely produce nonsense. They produce plausible words in place of the real ones. "Fifteen" becomes "fifty". A surname becomes a common word that sounds like it. "We did not agree" quietly loses its "not". Those errors read perfectly and change everything, which is why fluency is the thing to be suspicious of.
- Fix the load-bearing words first. Names, job titles, figures, dates and place names. These are the words a reader cannot reconstruct from context.
- Then read for negation and hedging. A dropped "not", "never" or "only" reverses a claim without disturbing the grammar.
- Then repair the sentence breaks. Machine punctuation splits long answers in odd places, and a misplaced full stop can change who a clause belongs to.
Getting speaker labels right
Interviews with more than two voices need one extra pass. Automatic speaker labelling, known as diarisation, is good but not infallible, and a misattributed answer is worse than a mistyped one. Anchor each label at the first moment a speaker is unmistakable: a name used, or a question only the interviewer would ask. Then read forward, checking every point where the conversation speeds up or people talk over each other.
Checking your own accuracy
For high-stakes work, sample five minutes against the audio and count the errors you find. That count converts into a rate you can compare between tools and over time; our guide to scoring transcription accuracy properly explains how. Verify every quotation you intend to publish against the audio itself, never against your memory of it.
Tip: Do the final read-through the same day if you can. You are the only person who can resolve an ambiguous passage from memory of the room, and that memory fades within days.
Formatting that keeps a transcript usable
A transcript is a working document, and formatting is what makes it work. The conventions below are close to universal. They answer the three questions every later reader asks: who said this, when, and how sure are we?
- Speaker labels. Use roles or initials — INT: and RES:, or surnames — applied identically throughout the file.
- Timestamps. Insert them at each speaker turn or every few minutes, written as hh:mm:ss. Any sentence in the text should be findable in the audio within seconds.
- Uncertainty markers. Write [inaudible 00:23:41] or [crosstalk] instead of guessing. A visible gap is a far smaller flaw than an invented word.
- Non-verbal cues. Note [laughs], [long pause] or [interrupts] only where they change the meaning. A dry answer and an ironic one can share every word.
- House style. Settle numbers, capitalisation, hyphenation and spellings once, then apply the same rules everywhere.
- Anonymisation. For research work, apply pseudonyms and strip identifying detail as your ethics protocol requires.
Save the finished transcript beside its audio, under a filename carrying the date and the participant. If the interview is heading for a podcast or a video, the same text becomes the raw material for captions. That conversion has its own rules, covered in our practical guide to SRT and VTT files. The full chain of stages, from consent form to archive, is indexed on our transcription workflow topic page.
Fixing the five common failures
Ruined transcripts fail in a small number of ways. Each has a fix, and most fixes belong earlier in the process than people expect.
| Problem | Usual cause | What to do |
|---|---|---|
| Whole passages unreadable | Microphone too far away; echoing room | Mark [inaudible] with timestamps, and move the microphone closer next time |
| Names and terms wrong throughout | No vocabulary list supplied to the engine | Search and replace once, then supply the list before the next interview |
| Answers given to the wrong person | Similar voices, crosstalk, one shared microphone | Re-anchor labels at unambiguous moments; record separate tracks in future |
| Reads well but misquotes | Fluent machine substitutions left unchecked | Verify every quotation against the audio before publication |
| Recording lost or corrupted | One device, no backup, battery failure | Always run a second recorder, and copy files off the device the same day |
Four of those five failures are prevented at the recording stage rather than the editing stage. That is the pattern worth internalising: transcription rewards preparation more than it rewards effort.
What the transcript is worth later
The immediate article or report is rarely the whole return. A transcript is searchable, so a half-remembered phrase from an interview two years ago takes seconds to find. It supports quotation with confidence, because every line carries a route back to the audio. It also makes the material accessible. The W3C's accessibility guidance treats transcripts as a core requirement rather than a courtesy. Deaf and hard-of-hearing readers depend on them to reach spoken content at all.
Archived properly, a body of interviews becomes an instrument instead of a pile of files. Researchers can code it for themes and reporters can revisit it. The whole set can also be searched at once, a task that AI assistants handle well in research and study work. That only holds if the underlying text is trustworthy. That proviso is the whole point of the correction pass.
Key takeaways
- Recording quality decides transcript quality: close the distance, choose a soft room, test for 30 seconds, run a backup.
- Ask for consent on the recording itself, and note the spelling of names and terms while the interview happens.
- Pick full verbatim, intelligent verbatim or summary notes before you start, then apply that convention consistently.
- AI drafts changed the economics, but the correction pass against the audio is where a transcript earns its trust.
- Speaker labels, timestamps and honest [inaudible] markers keep a transcript usable years after the conversation.
The tools will keep getting faster and the drafts will keep getting cleaner. The judgement in these steps does not automate. We make that argument more generally in why craftsmanship still matters in the age of AI. Everything else in the Voice & Transcription hub rests on the same division of labour. The machine supplies the speed. The care supplies the trust.