Editable audio transcription

Audio to Text

Turn a recording into an editable transcript, replay each time-marked section when a detail needs checking, then download the result as text, subtitles, or a PDF.

Upload audio to transcribeDrop a file here or browse from your device. MP3, WAV, M4A, AAC, OGG, WebM, MP4, MPEG, and FLAC are supported up to 25MB.

Editable transcript

Click a timestamp to replay that point in the audio, then edit the wording directly.

Your transcript will appear here

Upload an audio file, choose Transcribe, then review and edit the result beside the audio player.

From recording to a transcript you can actually check

Audio files are useful until you need to find a quote, turn a voice note into an email, make captions, or prepare meeting notes. Scrubbing through a long recording to locate one sentence takes time, while automatic transcription can mishear names, dates, prices, and specialist terms.

Audio to Text gives you an editable working transcript rather than a locked result. When timestamp data is available, each section is tied to its position in the player. Click the timestamp, hear the source again, and correct the words in place before exporting.

The optional custom-terms field helps with details that speech recognition can easily mishear: a person's name, company, product, location, acronym, or technical term. Keep the list short and focused on words that matter in this recording.

1. Upload audio

Choose an audio recording or a supported video file with spoken audio. The original file stays available in the player while you work.

2. Add context

Select the spoken language when you know it and add concise custom terms for important names or unusual words.

3. Review and edit

Replay a marked section from the transcript, then adjust its wording before any download is created.

4. Export the format you need

Copy the text, save a TXT file, create subtitles in SRT or VTT, or download a simple timestamped PDF.

Useful for everyday recordings

  • Voice notes: turn a spoken reminder or update into searchable text.
  • Lectures and interviews: create a reviewable first transcript, then replay uncertain sections.
  • Podcasts and video clips: prepare caption files or a draft description from spoken content.
  • Meetings: turn a recording into editable notes before organising it into a final document.

Important limits to understand

Speech recognition is not a substitute for checking important material. Overlapping voices, background sounds, heavy accents, low recording quality, and specialised vocabulary can cause mistakes.

The transcript editor is designed to make that review practical. Use the player to confirm names, amounts, dates, and statements before sharing or publishing the result.

For a polished document, download the transcript as a PDF or move the text into the Text to PDF editor to add layouts, headings, page settings, and other document controls.

Privacy and review

Your recording is used to produce the transcript shown in this tool and is not turned into a public link.

Automatic speech recognition can mishear names, figures, dates, overlapping speakers, or specialist language. Use the player and timestamps to check important passages before you share or publish the transcript.

Frequently asked questions

Which audio formats can I transcribe?

The tool accepts MP3, WAV, M4A, AAC, OGG, WebM, MP4, MPEG, and FLAC files up to 25MB. MP4 and WebM are useful when the spoken audio is inside a video recording.

Can I correct the transcript before downloading it?

Yes. Every returned section is editable. Click its timestamp to hear that part of the source audio, then correct a name, date, number, or sentence before exporting.

Can I create subtitles from an audio file?

Yes. Download SRT for common subtitle workflows or VTT for web captions. The timing comes from the transcription response when the selected transcription model provides segment timings.

Will Audio to Text identify different speakers?

The initial tool focuses on a reliable editable transcript and time-marked segments. Speaker identification is a separate higher-compute workflow and is not automatically added to every file.

Can I add names or specialist terms?

Yes. Add a short list of names, brands, locations, acronyms, or specialist words that may be spoken. This gives the transcription extra context for terms that are easy to mishear.

Is an audio transcript always correct?

No. Background noise, overlapping speakers, accents, poor microphone quality, and unfamiliar names can affect automatic speech recognition. Replay important sections before relying on a transcript for legal, financial, medical, or published material.

What happens to the uploaded audio?

The recording is used to create your transcript and is not published as a public file. Review the result before sharing it, especially when the recording contains sensitive information.