The Complete Guide to Interview and Meeting Transcription: From Recording to SRT/VTT Export
The Complete Guide to Interview and Meeting Transcription
The real work starts after the interview ends. Manually transcribing one hour of interview audio takes 4–5 hours; one meeting carries 20,000+ characters of information, and summarizing it still eats "someone's evening" in most workplaces.
This guide walks the full practical workflow: recording tips → AI transcription → editing → export to minutes/drafts/subtitles. I'll use PUBVOICE as the example tool, but most of it — recording principles, SRT vs. VTT — applies to any transcription service.
The Reality of Time and Cost
| Method | Rough time | Notes |
|---|---|---|
| Manual transcription | 4–5 hrs per audio hour | Includes relistening |
| Human outsourcing | Days of turnaround | ~thousands of yen per hour |
| AI + editing | 10–30 min | Fixing misrecognitions, verifying |
What AI replaces is typing time, not verification time. But removing typing alone shrinks the whole job to under a tenth — that's the premise for designing the workflow.
Step 0: Recording — 80% of Quality Is Decided Here
The biggest factor in transcription accuracy is the recording, not the AI:
- Get mics close to speakers. One phone picking up a whole room is the worst case for diarization. Put a recorder near each speaker, or arrange seating so everyone's voice lands clearly.
- Reduce crosstalk. Simultaneous speech wrecks both recognition and speaker separation. A facilitator saying "one at a time, please" saves hours of editing later.
- Use standard formats (mp3/m4a/wav). PUBVOICE accepts files up to 500MB and recordings over 5 hours.
- Record a 30-second test and check that the quietest person is audible. This alone prevents most failures.
Step 1: Upload and Transcribe
In PUBVOICE: drag the file into the workspace (/transcripts), watch status update (queued → processing → done), and get a transcript with speaker labels and timestamps per segment. New users can try up to 90 seconds of audio on the top page without registering — test it with the kind of audio you actually record.
Step 2: Edit — Fix "Who Said What"
AI misrecognitions fall into three practical buckets:
- Proper nouns and jargon (company names, people, industry terms) — the most common
- Speaker assignment swaps — when speakers sit close or sound alike
- Homophone choices — fix by context
The basic flow in PUBVOICE's editor: bulk-rename speakers (Speaker 1 → "Tanaka"), then use in-document search to hunt down misrecognitions. The editor highlights segments in sync with the audio player, so "click a suspicious line, relisten, fix" happens in one place — you verify while the conversation is still fresh, which is the real advantage over outsourcing.
Step 3A: Turn It into Minutes
The key to minutes is not writing everything:
- Pull decisions out first (that's the body of the minutes)
- Table the TODOs (who, by when)
- Quote background only where the trail matters, briefly, from the relevant segments
- Close with open items and next agenda
Keep the timestamped transcript as the traceable archive; share the TXT (speaker-labeled) export as the base document.
Step 3B: Turn It into an Article
- Rename speakers to real names, split into readable paragraphs
- Mark the quotes you'll use — same as in the paper era
- Order quotes to your structure and write the connective text
- Verify numbers and proper nouns against the audio by clicking the timestamp
That last step matters most: numeric misquotes kill articles, and one-click relistening collapses the fact-check loop.
Step 3C: Subtitles — SRT vs. VTT
| Format | Main use | Notes |
|---|---|---|
| SRT | YouTube, Premiere, Final Cut, most editors | Most widely supported. HH:MM:SS,mmm (comma) |
| VTT | HTML5 <video>, web players |
Needs WEBVTT header. HH:MM:SS.mmm (period). Styleable |
| TXT (speaker-labeled) | Minutes, drafts, archives | Readable text with speakers and times |
When in doubt, choose SRT — nearly every video tool reads it. For direct web embedding, use VTT. PUBVOICE exports all three from a single transcript.
Security and Handling Notes
- Consent to record: legal in Japan in principle, but telling interviewees and participants that you're recording — and why — is the professional standard (and mandatory in some industries/regions)
- Confidential material: for board/legal/HR audio, check the service's data location and retention first; mask sensitive parts during editing when external transmission is a concern
- Retention design: when transcripts serve as records, decide the retention period for source audio and transcripts up front
FAQ
Q. Does it handle long recordings (2–3 hour lectures)?
A. Yes — files up to 500MB, with speaker separation that keeps Q&A sections attributed.
Q. How about dialects and non-standard speech?
A. Standard Japanese yields the best accuracy; dialects, fast speech, and mumbling increase errors — which is why Step 0 matters most there.
Q. What about English and other languages?
A. Processable, but current optimization is Japanese-first. If multilingual is your main use case, consider multilingual services like Notta from the comparison article.
Q. Can I share results with my team?
A. Exported TXT/SRT/VTT files share as-is. For minutes operations, the standard pattern is a summary (decisions + TODOs) for circulation plus the timestamped full text as the record.
Related articles

Yutaro Sasao
CEO / MediaLeap Inc.
After leading web media monetization and data analytics at KADOKAWA / DWANGO, and driving programmatic ad revenue growth in SSP / ad network businesses, he founded MediaLeap Inc. in May 2025. He now develops and operates AI audio SaaS "PUBVOICE", tourism DX app "ANIME TRAVEL", and AI voice chat app "AITOMO". Drawing on cross-functional expertise in advertising, technology, analytics, and business, he works to improve media revenue through data-driven strategies.
Related posts
AI Transcription Services Compared (2026): A Practical Guide for Individuals and Teams
Notta, toruno, Rimo Voice, Otolio, AI GIJIROKU, YOMEL, and PUBVOICE compared on free tiers, effective per-hour pricing, diarization, editing, and export formats — with task-by-task recommendations and annual cost examples.
AI Transcription Market: A Sell-Side Analyst's Breakdown (2026)
The speech recognition market grows from $9.66B (2025) to a projected $23.11B by 2030 (~19% CAGR) while API costs fell ~10x in three years. We break down market sizing, the cost curve, Japan demand data, competitive structure, and unit economics at sell-side research depth.