Short answer: AI transcription is only as private as the provider's policies and your own habits. Before you upload, check three things: whether your audio trains their models, how long files are kept, and which country stores them. Then guard the transcript itself, because text spreads much faster than audio ever does.

A voice recording carries far more than words. It holds the sound of one specific person, which is an identifier in its own right. It holds names, diagnoses, salaries and half-formed ideas. It also holds the voices of people who never agreed to be recorded. All of that travels with the file when you press upload.

AI transcription solved a real problem. Work that once took a typist four hours now takes a few minutes. The trade is that huge volumes of recorded speech now move through cloud pipelines. So the question has changed. It is no longer whether a machine can transcribe your audio accurately. It is what happens to the recording once it can.

This guide belongs to the Voice & Transcription hub. It maps what a recording actually contains, where the risk sits, which rules apply, and what to ask a provider first.

What is actually inside a voice file

An honest inventory helps, because the phrase "an audio file" undersells the contents. Five distinct kinds of data ride along in a single recording, and only one of them ends up visible in the transcript.

  • The voice itself: a voice can identify a speaker. Under the GDPR, a voiceprint used to recognise someone counts as biometric data, a special category with stricter rules attached.
  • The spoken content: people say things aloud that they would never write down. Diagnoses, salaries, disputes and private doubts all surface in ordinary speech.
  • The bystanders: recordings catch people who never consented. The colleague passing the desk and the child in the next room both land in the file.
  • The metadata: filenames, timestamps, calendar entries and device details show who spoke to whom and for how long. Patterns of contact reveal a great deal on their own.
  • The transcript: once speech becomes text it becomes searchable and easy to copy. Transcription does not create sensitivity. It makes sensitive material far easier to find and forward.

Follow one file from microphone to deletion

Risk is easier to manage once you can see where it sits. Trace a single recording through its life and the exposure clusters at five points.

  1. Capture: the questions start before any software runs. Does everyone in the room know a recording is running, and did anyone actually agree?
  2. Transit: audio should cross the internet over an encrypted connection. The bigger risk is informal transit, such as raw files emailed as attachments or dropped into a chat thread.
  3. Processing: your audio meets the models that perform the conversion. Ask whether it is kept afterwards, and whether any human reviewer listens to samples for quality control.
  4. Storage: defaults cause most of the damage here. Many services keep uploads and transcripts indefinitely unless you tell them not to.
  5. Afterlife: the finished text is the leakiest stage. It gets pasted into emails, tickets, minutes and slide decks, and every copy sits outside the original controls.
Your audio is only as private as the least protected copy of its transcript.

None of these stages is exotic. Each one simply has a default, and defaults favour convenience over caution. Most of the work here is noticing them. Decide who gets told, what gets kept and where it goes, then change the few settings that matter for your material.

Processing is the stage with no equivalent in older workflows. A typist heard your tape and typed it. A model turns sound into probabilities, and that pipeline may involve several companies. Our explainer on what really happens to audio inside a recogniser shows which steps are involved.

What the security words on the page actually mean

Vendor security pages lean on a small vocabulary. Knowing what each term promises, and what it does not, makes those pages far easier to judge.

Encryption in transit and at rest

In transit means the audio is scrambled while it crosses the internet, normally using TLS. At rest means the stored file is scrambled on the provider's disks, usually with AES-256. Both are baseline expectations rather than features. Neither one stops the provider reading your file, because the provider holds the keys.

TLS 1.3current protocol standard for encrypting audio in transit
AES-256common cipher strength for encrypting stored files at rest
72 hoursGDPR deadline for reporting a personal data breach to a supervisory authority

Retention, deletion and zero data retention

Retention is how long a service keeps your files. Zero data retention means the audio is discarded as soon as the transcript is returned. Ask what deletion actually covers. A good answer includes backups and names a window, such as thirty days, instead of promising something vague.

On-device, private cloud and shared cloud

Some speech models now run on the device that made the recording, so no audio leaves it. Others run in a dedicated cloud instance, and most run on shared infrastructure. On-device work is the strongest privacy position available. It usually costs some accuracy on difficult audio, which makes it a real trade rather than a free win.

Tip: Ask for the subprocessor list, not just the security page. Most transcription products are assembled from other companies' services, and that list tells you who else can touch your audio.

Which rules apply to your recordings

No single law governs transcription. Three families of rules cover most situations, and one set of voluntary standards is worth knowing. None of what follows is legal advice.

FrameworkWhat it coversWho it bindsWhat it asks of you
GDPR and UK GDPRRecordings and transcripts of identifiable peopleAny organisation processing that data, wherever basedA lawful basis, purpose limits, data minimisation, access and erasure rights
HIPAAProtected health information in US clinical settingsCovered entities and their business associatesA signed business associate agreement plus technical and administrative safeguards
Consent-to-record lawsThe act of recording a conversation at allAnyone recording a call, meeting or interviewOne-party or all-party consent, depending on the state or country
ISO/IEC 27001 and SOC 2Voluntary security standards, not lawVendors that choose to be auditedEvidence that practices were examined, not merely described

The GDPR treats a voice as personal data

The EU's General Data Protection Regulation covers recordings of identifiable people. That brings the full set of duties into play. For most teams the practical reading is short. Record for a stated reason, tell people, keep less than you think you need, and know where your vendor stores it.

HIPAA turns it into a contract question

In the United States, clinical dictation and patient calls fall under HIPAA for covered organisations. The practical effect is contractual. A vendor that handles protected health information must sign a business associate agreement. A consumer tool without one is not an option, however accurate its output.

Consent laws vary more than people expect

Data protection and recording consent are separate questions. Some places require consent from only one party to the conversation. Others, including California, require consent from everyone. One dull habit satisfies nearly all of them. Announce the recording at the start, say why, and let that announcement appear in the transcript.

Decide the sensitivity tier before you upload

Not every recording needs the same care, and treating them all alike wastes effort in one direction or invites trouble in the other. Sorting audio into three tiers takes seconds.

  • Tier one, routine: internal stand-ups, public talks, your own voice notes. A mainstream cloud service with training turned off and a short retention window is fine.
  • Tier two, confidential: client calls, hiring interviews, unreleased plans. Require a named retention period, no training on your audio, and controlled access to the finished transcript.
  • Tier three, regulated or high risk: health data, legal matters, journalism with vulnerable sources. Require a contract that names the obligations, audited certification, and regional or on-device processing where possible.

The tier decides the tooling, not the other way round. Defending one strict workflow for tier three is straightforward. Arguing after the fact about a file that should never have gone through a free web form is not.

Four myths worth retiring

  • "Deleting the file deletes the data." Copies persist in backups, caches and shared drives. Deletion only counts when the provider states which systems it reaches, and by when.
  • "Redaction makes a transcript anonymous." Removing names is rarely enough. Job titles, dates, places and distinctive turns of phrase let a motivated reader reassemble who said what.
  • "Local processing is automatically safer." A laptop full of unencrypted recordings with no retention rule is not safe. On-device work only helps when the device itself is managed.
  • "The audio is the risky part." The transcript usually leaks first. It is small, searchable and easy to paste, so it travels much further than the original file.

The redaction myth deserves extra care, because it is easy to over-trust. Redaction lowers exposure. It does not remove it. Access control and retention limits still matter for redacted text, which is why these habits work as a set rather than a menu.

Eight questions to ask before you upload

Security pages are written to reassure. These eight questions get past the reassurance quickly, and they are worth asking in writing.

  • Is audio encrypted in transit and at rest? Expect named protocols and cipher strengths, stated plainly rather than implied.
  • Is my audio used to train your models? This is the most important question on the list. Look for a clear default and an opt-out you do not have to negotiate.
  • How long do you keep audio and transcripts? A good answer names a period, and lets you shorten it or delete on demand.
  • Who inside your company can access my files? Access should be role-limited, logged and rare, never routine.
  • Do human reviewers ever listen to samples? Quality review is legitimate work, but you should know it happens and be able to decline.
  • Which independent audits do you hold? A certification such as ISO/IEC 27001 or a SOC 2 report shows the practices were tested by an outside party. Verified quality beats claimed quality, a principle we cover in why an audited figure outranks a stated one.
  • Where is my data processed and stored? Data residency matters for GDPR transfers and for many sector-specific rules.
  • What happens if you are breached? Look for a named process and a notification deadline, not aspirations.

Free tools deserve one extra question: what pays for this service? Sometimes the answer is a fair freemium model. Sometimes the answer is your data. We look at what free and paid AI tools really give you in more detail.

Habits that keep transcripts safe

Vendor diligence is half the job. The rest is behaviour inside your own team, and it costs little more than deciding to have rules at all. Formal frameworks such as the NIST Cybersecurity Framework organise this thinking at company scale. The essentials fit in a short list.

  • Announce, then record. One sentence at the start prevents an argument later. Our guide to running meeting transcription well sets out the wording.
  • Record less. The safest recording is the one you never made. "We might want it someday" is how archives of liability accumulate.
  • Separate the sensitive files. An interview with a named source needs tighter handling than a status call. Our guide to recording an interview responsibly shows where consent fits into the process.
  • Store transcripts like documents. Keep them where permissions and retention policies already exist, not in chat threads where they stay findable forever.
  • Set a retention schedule and honour it. Decide how long transcripts live, then delete on time. Text you no longer hold cannot leak.
  • Extend the rules to AI chat tools. Pasting a transcript into a chatbot is a disclosure like any other. Our guide to responsible AI habits covers the same ground.

Tip: Put the retention rule in the file name. A transcript called "2026-03-board-call-delete-2026-09" tells the next person what to do without anyone reading a policy document.

The confidentiality question is not only about software. Handing a recording to a vetted professional narrows exposure to one named person bound by a contract you can enforce. Sending it to a service widens the circle to a company and its subprocessors, yet often means nobody listens at all. Neither shape is safer in the abstract, which is why we set them against each other in which method deserves your confidential audio.

The encouraging news is that the technology is drifting towards privacy. Speech models small enough to run on a phone, automatic redaction of names and numbers, and short default retention are becoming ordinary rather than premium. The same expectations shaped our own TRANSCRIPT.YOU tool. What will not change is the judgement involved. A transcription service is a custodian of conversations, and it deserves the scrutiny you give anyone you trust with something valuable. Further reading across the journal is collected on our AI privacy topic page.

Key takeaways

  • A recording carries identity, candid content, bystanders and metadata, not just the words on the page.
  • Risk clusters at five points: capture, transit, processing, storage and the transcript's afterlife.
  • Encryption in transit and at rest is a baseline; retention limits and training opt-outs are what actually vary.
  • The GDPR, HIPAA and consent-to-record laws apply differently, but announcing the recording satisfies most of them at once.
  • Sort recordings into tiers, then ask providers in writing about training, retention, audits and data residency.