Tools
Turn a .vtt transcript into clean text
Trippi Notes team · facts checked on 2026-10-01
A .vtt file is a caption file. Every few seconds of speech gets an id, a time range and a short line, so a one-hour call turns into a thousand fragments. This converter turns it back into who said what, ready to read, quote or paste into an AI chat.
Open a file or paste a transcript. Nothing is uploaded: the file is read by this page only.
The file never leaves your browser. The converter is a short script inside this page. It reads the file on your computer and writes the result into the box above. There is no upload, no account and no copy kept anywhere. Load the page, switch off your network, and it still works.
Where the .vtt file comes from
Microsoft Teams. After a meeting with transcription, open the meeting chat, choose Recap, then Transcript, and pick the file type next to Download. Microsoft says that by default the organizer and co-organizers can download it, as a .docx or a .vtt file; your IT admin can let others do it too. Take the .vtt for this tool.
Zoom. Zoom writes an audio transcript for meetings recorded to the cloud, and saves it in VTT format. It needs a Pro, Business, Education or Enterprise account, cloud recording, and audio transcription switched on in the recording settings. The file appears next to the recording once it is processed.
Anything else that writes WebVTT works the same way: caption files from video editors, course platforms or video sites. Old .srt files mostly work too, because their timing lines look almost the same.
What the converter changes
Here is the start of a typical Teams file:
WEBVTT 3f1c2a0e-1 00:00:04.120 --> 00:00:07.480 <v Anna Berg>Okay, let's start with the launch date.</v> 3f1c2a0e-2 00:00:07.480 --> 00:00:10.900 <v Anna Berg>We are moving it to October 3rd.</v> 3f1c2a0e-3 00:00:11.200 --> 00:00:13.050 <v Jonas Weber>Then I need the copy by Friday.</v>
And this is what comes out with timestamps kept and lines merged:
[00:00:04] Anna Berg: Okay, let's start with the launch date. We are moving it to October 3rd. [00:00:11] Jonas Weber: Then I need the copy by Friday.
- Header and notes are dropped. The WEBVTT line, NOTE and STYLE blocks never carry speech.
- Cue ids go. Teams numbers each cue with a long id like 3f1c2a0e-1. Nobody reads those.
- Timestamps become short or go. A timing line such as 00:00:04.120 --> 00:00:07.480 becomes [00:00:04], or nothing if you untick Keep timestamps.
- Speaker names are kept. Teams puts the name in a voice tag, <v Anna Berg>, which is part of the WebVTT standard. Zoom files usually start each line with the name and a colon. Both become “Name: text”.
- Other tags and codes are removed. Italic or class tags and HTML codes such as & turn into plain characters.
- Fragments are joined. With Merge on, everything one person says in a row becomes one paragraph. Exact repeats in a row are dropped.
Keep timestamps or not?
Keep them when someone will check the text against the recording, or when you want to quote a moment: “at 00:41:10 Jonas agreed to the date”. They also help when a transcript is long and people refer to “the part near the end”.
Drop them when the text is for reading, or when you paste it into ChatGPT or Claude. A timestamp on every line adds about a fifth to the length and gives the model nothing to work with. Shorter input also means a long call is more likely to fit into one message.
Plain text or Markdown
.txt is the safe choice. It opens everywhere and pastes cleanly into an email, a ticket or an AI chat.
.md is Markdown: the file gets a heading and every speaker name is set in bold. Pick it for a wiki, a notes app or a repository that shows Markdown, where bold names make a long talk easier to scan.
If you need a Word document, copy the text and paste it into Word. Teams can also give you a .docx directly from the same Download menu.
When the result looks wrong
- No names at all. Names come from the file. If the tool that wrote it did not name speakers, there is nothing to keep. In Zoom you can fix names on the recording page before you download.
- A word ended up as a name. A line like “Plan: we ship Friday” looks the same as “Anna: we ship Friday”. The converter only treats short phrases of up to five words before a colon as a name; rename the rare miss by hand.
- Lines repeat in steps. Some auto-caption files repeat the last words of a line at the start of the next one. Exact repeats are removed, partial ones are not.
- Strange characters. The file should be UTF-8. If accented letters look broken, open the file in a text editor, save it as UTF-8 and try again.
After the text: minutes and a follow-up
Clean text is the raw material, not the result. To turn it into minutes, paste it into the transcript to minutes prompt builder, which prepares a prompt for ChatGPT or Claude. For the structure of good minutes, see how to write meeting minutes; for working with a chat model, ChatGPT for meeting notes.
Or skip the .vtt step
Trippi Notes is a Chrome extension that writes the transcript while the call is happening in a browser tab: Google Meet, Zoom in the browser or Microsoft Teams in the browser. Each line has the speaker’s name and the time. When the call ends you export it as Markdown or plain text, with no admin setting, no cloud recording and no file to download.
It starts only when you click its icon in the meeting tab, and it hears the tab, not the meeting chat. No bot joins the call, so tell people you are taking notes; here is how to say it. The desktop Zoom and Teams apps are out of its reach.
FAQ
Asked about this.
Is my file uploaded anywhere?+
No. The page reads the file with your browser and converts it with a script that runs on your computer. Nothing is sent to us or anyone else, and nothing is stored after you close the tab.
How do I get the .vtt file from Teams?+
Open the meeting chat, choose Recap, then Transcript, and choose .vtt next to Download. By default only the organizer and co-organizers can do this; Microsoft’s guide explains the rest.
Why does my result have no speaker names?+
Because the file has none. The converter keeps names from Teams voice tags and from Zoom “Name:” lines, and does not guess names that are not there.
Does it work with .srt subtitle files?+
Mostly yes. SRT uses numbered blocks and timing lines with a comma before the milliseconds, and the converter reads those too. SRT has no speaker tags, so names appear only if they are written in the text.
Can I use it on a phone?+
Yes, if your phone lets you pick the file. You can also open the file, copy its text and paste it into the box.
Read next
Get the transcript without the file.
Trippi Notes is in the Chrome Web Store. One click on the icon when the call begins, and nothing before that.
