How to Transcribe Audio for Free: Google Docs Voice Typing, Gemini, Whisper, and Your Phone's Built-In Features — What They Can Do, Where They Stop, and How to Choose
Transcribing Audio for Free: Four Methods Compared
The author of this article is Yutaro Sasao. He develops and runs PUBVOICE, an AI voice SaaS with paid plans, but this article evaluates the free methods fairly. Free tools cover more than you might expect — whether they are enough depends on your audio and how you use it.
Meeting minutes, interviews, lecture recordings, video subtitles — transcription needs keep growing, but retyping by hand is not realistic. According to a Ricoh column, manual transcription generally takes four times the length of the audio Ricoh.
The conclusion first. Choosing a free method comes down to four lines:
- Transcribe one person's speech right now → Google Docs voice typing, or your phone's built-in dictation (both completely free)
- Transcribe a short recorded file, with a summary too → Gemini (free tier)
- Process long or bulk audio, repeatedly → running OpenAI Whisper yourself (free, open source)
- Need speaker separation for multi-person conversations → beyond these methods; a free tier of a speaker-separating service is the realistic answer
Note that "free" comes in two kinds: completely free methods (voice typing, phone features, open source) and services with a free tier (Gemini, tools like PUBVOICE). The author selected PUBVOICE's speech recognition and voice generation engines himself, cut development costs to roughly a third with generative AI, and also develops AITOMO, a VOICEVOX-based voice AI app with 20,000 downloads and a 4.2 rating — so what follows is written from hands-on experience with both cloud models and open-source engines.
Method 1: Google Docs Voice Typing
Open a document in Chrome, choose "Tools" → "Voice typing", click the mic, and your speech is transcribed in real time. It costs nothing with a Google account; the official help lists the latest Chrome, Edge, and Safari as supported browsers and over 100 languages Google.
Strengths: minimal setup and freedom of length — nothing to install, no practical limit while you keep talking. Ideal for drafting by dictation.
Three limits: there is no feature to import an existing audio file (voice typing listens to your microphone, so an MP3 requires routing speaker output back into the mic, which is unreliable); no speaker separation, so multi-person conversations merge into undifferentiated text; and voice commands for editing and formatting are English-only per the help Google, so in Japanese treat it as a "writing" tool and fix errors with the keyboard.
Method 2: Your Phone's Built-In Voice Features
Before searching for a free transcription app, check what your phone already has.
On iPhone, enable dictation under Settings → General → Keyboard → Enable Dictation, then tap the mic button on any keyboard Apple. Many languages are processed on the device without internet; supported languages get automatic punctuation; and dictation stops after 30 seconds of silence Apple. Android ships with voice input in Gboard; advanced features have language and device conditions (such as Pixel 6 and later), so check the official help Google.
Also notable on iPhone: the stock Notes app can record audio and show a live transcript while recording, saving both into the note Apple. Japanese is among the ten supported languages, requiring iPhone 12 or later Apple.
The shared limit is real-time recognition: they do not transcribe existing files (Notes only transcribes recordings made inside Notes), and there is no speaker separation. For transcribing recorded files later, see the voice memo transcription guide.
Method 3: Handing Audio Files to Gemini
In the Gemini app, upload an audio file with the attach button and prompt "transcribe this audio". Per the official help, audio is limited to a total of 10 minutes on the free plan and 3 hours on paid plans (Google AI Pro/Ultra), with up to 10 files per prompt (as of October 2026) Gemini.
The strength is that transcription is part of a generative AI: "transcribe, then summarize in three points" or "format as minutes" can be requested in one instruction. It suits personal study and light meeting memos.
Four cautions: length (the free tier caps audio at 10 minutes, so an hour-long meeting does not fit); speaker separation is not guaranteed (the help does not document it); accuracy attitude (generative AI can fill inaudible parts with plausible text, so never use the output as quotes or records without checking against the original audio); and confidentiality (you are handing audio to an external service, so check the official pages on storage and use for AI improvement — free and paid plans may differ).
Method 4: Running OpenAI Whisper Yourself
Whisper is an open-source speech recognition model from OpenAI, released under the MIT license and free to run on your own PC. One model handles multilingual recognition, translation, and language identification, Japanese included GitHub.
You need a Python environment, PyTorch, and basic command-line skills. Six model sizes exist, with GPU memory requirements from about 1GB (small models) to about 10GB (large) GitHub. Computers without a GPU can run it on CPU, but processing time grows sharply for long audio. The basic step — install with pip, run one command, get text — is not hard; the real effort comes after, because speaker separation is not part of Whisper itself, and you must build the surrounding pieces: processing plans for long files, a playback UI for checking errors, and output management.
From the developer's first-hand view: the author selected PUBVOICE's speech recognition engines and evaluated both cloud APIs and open-source models. On accuracy alone, Whisper-class models are fully practical. But transcription as a service is built around the recognition — speaker separation, an audio-synced editor, export formats — and that surrounding work is exactly what the author has shouldered while developing AITOMO. For confidential audio that must stay local, large volumes, and people comfortable with tooling, Whisper remains the free option of choice.
Comparison Table
| Method | Cost | Speaker separation | Length guide (as of writing) | Skills needed | Best for |
|---|---|---|---|---|---|
| Google Docs voice typing | Completely free | None | Effectively unlimited while speaking | None | One-person dictation, drafting |
| Phone built-ins (iPhone dictation, Notes recording, etc.) | Completely free | None | Real-time focus (Notes transcribes recordings made in Notes) | None | Mobile memos, short dictation |
| Gemini | Free tier | None (not guaranteed) | 10 min free / 3 h paid for audio | None | Short audio + summarizing |
| Self-run Whisper | Completely free (OSS) | Not built in (extra libraries) | As much as your PC handles | Python, CLI | Long, bulk, confidential audio |
Google, Apple, Gemini, and OpenAI specifications reflect official information as of October 2026 (Google Docs help, Apple guide, Gemini help, Whisper repository). Free-tier minutes and supported languages change, so confirm on the official pages before use.
When Free Is Enough, and When to Pay
- Speaker separation. For one speaker, free methods suffice; the moment "who said what" matters, these four fall short — the clearest trigger to move to a service. PUBVOICE's free plan offers 60 minutes of speaker-separated transcription monthly, no credit card required
- Volume and frequency. A few short files a month fit the free toolkit; a weekly one-hour meeting argues for starting on a free tier and scaling with volume
- Confidentiality. Audio that cannot leave your machine belongs with Whisper. For cloud services, check storage and AI-training use first. PUBVOICE never uses uploaded audio for AI training
- Accuracy requirements. For quotes and records of decisions, verify against the original audio in an audio-synced editor; never ship generative AI transcripts unverified
- Continuity. One-off needs should stay free; team-wide use should be judged on sharing and formatting too
If unsure, run one free method on your own audio, then pay based on the limit you actually hit. For service comparisons, see the 2026 AI transcription comparison.
FAQ
Q. Can Google Docs voice typing transcribe an existing MP3 file?
A. No. Voice typing listens to your microphone in real time and has no file import. For recorded files, use Gemini (if short), Whisper, or a file-capable service — see the voice memo transcription guide.
Q. Can free AI transcription do speaker separation?
A. Not with the four methods here. Start with a free tier of a service that supports it; PUBVOICE's free plan includes 60 minutes of speaker-separated transcription monthly.
Q. May I upload company meeting audio to Gemini?
A. For confidential meetings, check the official guidance on storage and AI improvement use first. If confidentiality is the priority, Whisper running locally is the candidate.
Q. Does Whisper run on phones or ordinary laptops?
A. It runs on CPU, but processing time grows sharply for long audio. Comfortable use assumes a PC meeting the repository's GPU memory requirements, roughly 1–10GB depending on the model GitHub.
Q. How much can I transcribe for free?
A. Real-time voice typing is free for whatever you speak; Gemini's free plan covers 10 minutes of audio total (as of writing); Whisper is free as long as your PC can process it. PUBVOICE's free tier offers 60 minutes of transcription and 10,000 characters of voice generation monthly, no credit card.
Related Articles
- Meeting Minutes: The Complete Guide — a copy-paste template and an AI transcription workflow for minutes
- Interview and Meeting Transcription: The Complete Guide — the full workflow from recording to SRT/VTT export
- AI Transcription Services Compared (2026 Edition) — pricing, free tiers, and export formats side by side
- iPhone Voice Memos and IC Recorder Transcription Guide — concrete steps for transcribing recorded files
- Text-to-Speech: Turning Text into AI Audio — how to turn scripts and documents into audio files, and how to choose a read-aloud tool
- Transcription Help Guide — step-by-step instructions inside PUBVOICE
Free methods are enough to start; choosing well keeps them working. Run one on your own audio first. When you need speaker separation and an audio-synced editor, PUBVOICE's free plan gives you 60 minutes of transcription and 10,000 characters of voice generation monthly, with no credit card — see the transcription help guide, and upload your next recording to PUBVOICE.

Yutaro Sasao
CEO / MediaLeap Inc.
After leading web media monetization and data analytics at KADOKAWA / DWANGO, and driving programmatic ad revenue growth in SSP / ad network businesses, he founded MediaLeap Inc. in May 2025. He now develops and operates AI audio SaaS "PUBVOICE", tourism DX app "ANIME TRAVEL", and AI voice chat app "AITOMO". Drawing on cross-functional expertise in advertising, technology, analytics, and business, he works to improve media revenue through data-driven strategies.
Related posts
Transcribing Voice Memos and IC Recorder Audio: The Complete Guide from File Export to AI Transcription
A complete guide to transcribing audio from iPhone Voice Memos and IC recorders: recording tips, exporting m4a and mp3 files, and choosing among built-in phone features, self-hosted Whisper, and AI transcription services.
AI Transcription Services Compared (2026): A Practical Guide for Individuals and Teams
Notta, toruno, Rimo Voice, Otolio, AI GIJIROKU, YOMEL, and PUBVOICE compared on free tiers, effective per-hour pricing, diarization, editing, and export formats — with task-by-task recommendations and annual cost examples.
How to Read Text Aloud and How to Choose an AI Text-to-Speech Tool
A practical guide to turning the text you already have into audio: how to use the built-in read-aloud features of Windows, Mac, iPhone, and Android, how AI voice generation differs, and a checklist for choosing a read-aloud tool.