How to Read Text Aloud and How to Choose an AI Text-to-Speech Tool
How to Read Text Aloud: The Complete Guide to Choosing Your Method
The author of this article develops and runs PUBVOICE, an AI voice generation tool. The built-in read-aloud features and selection criteria described here do not depend on any specific tool and apply to whichever service you use.
Drafts, work documents, study materials, bedtime-story scripts — text you already have can be re-consumed by ear, on a commute or between chores. The conclusion first: there are three routes to read text aloud. Built-in OS and device read-aloud features, AI voice generation services, and dedicated read-aloud apps. Which one fits comes down to a single question: do you just need it read aloud on the spot, or do you need an audio file (MP3) you can take with you?
This article is for individuals who want to paste in their own text and get an audio file out. Turning articles on your own blog or media site into an embedded audio player is a different, publisher-facing workflow — that one is covered in The Complete Guide to Turning Blog Articles into Audio.
I have worked on voice content in the advertising and media industry since 2013, and today I develop PUBVOICE, an AI voice generation tool. This article builds the yardstick for choosing a tool first — including my own.
What This Article Covers
- The three routes to read text aloud, and what each is good and bad at
- How to use the built-in read-aloud features of Windows, Mac, iPhone, and Android, with official guides
- Where AI read-aloud differs: voice naturalness, emotional expression, and audio file export
- A 7-point checklist for choosing a read-aloud tool
- The step-by-step flow from pasting text to downloading an MP3 with PUBVOICE
Five Situations Where Read-Aloud Earns Its Keep
- Turning documents into audio. Internal reports and newsletters listened to on the commute.
- Study materials as listening practice. Memorization summaries and language material played on repeat — hands free, so it pairs well with walking or chores.
- Bedtime-story scripts. Streaming the script as audio before reading it aloud yourself, to check pacing and wording.
- Proofreading drafts. Sentences that look fine on screen reveal their clunkiness when heard aloud — useful for presentation scripts and narration.
- Resting your eyes. Long documents received by voice between hours of screen reading.
What these share: the text already exists. And in most of them, an audio file you can carry makes the workflow better — on a phone, shared, archived. Whether a tool can produce that file is exactly where built-in features and AI voice generation part ways.
The Three Routes to Read Text Aloud
| Route | Cost | What it does | Main limits | Best for |
|---|---|---|---|---|
| 1. Built-in OS/device read-aloud | Free | Reads the screen or selected text on the spot | Basically cannot save audio files; limited voices | When on-the-spot playback is enough |
| 2. AI voice generation services | Free tiers | Natural voices, exports MP3 and other formats | Per-generation character limits; advanced features vary by plan | When you want an audio file to carry |
| 3. Read-aloud apps and software | Free ones exist | Playback features specialized for files and books | Wide feature variance by purpose | Repeated listening to texts you keep |
The first fork is simply: do you need an audio file? If listening right now is enough, the built-in features carry it. If you want to listen on the move, edit, or share, routes 2 and 3 come in.
Free and Immediate: Built-in Read-Aloud Features
Every device ships with read-aloud built in — accessibility features originally, free for anyone, reversible anytime.
Windows. The Narrator screen reader reads on-screen elements and text; Microsoft publishes a chaptered complete guide Microsoft. For web pages, Edge's Immersive Reader is the easy route: enter it from the address bar or by right-clicking selected text, then have the page read aloud with adjustable text spacing and line focus Microsoft.
Mac and iPhone. On the Mac, turn on "Speak selection" under System Settings, Accessibility, then press Option+Esc (default) on selected text; speed, voice, and word highlighting are on the same screen Apple. On iPhone and iPad, enable "Speak Selection" and "Speak Screen" under Settings, Accessibility, Spoken Content Apple.
Android. TalkBack reads the whole screen and turns on from Accessibility settings or by voice Google. Select to Speak reads just what you tap or select — including text captured with the camera Google.
The shared limit. All of these read aloud on the spot — and that is all they do. Saving the reading as an audio file, carrying it on your phone, handing it to someone: not what built-in features are for.
Where AI Read-Aloud Differs: Naturalness, Emotion, and the File
Three differences matter. First, naturalness: built-in voices are flat and tiring over long passages, while AI voices breathe and stress like a human reader — PUBVOICE uses ElevenLabs as its voice engine. Second, emotional control: some AI tools accept tags embedded in the text to switch emotion and tone mid-read. Third, and the biggest in practice: generated audio can be downloaded as an MP3. The reading becomes a file you carry, share, and archive. On top of that, you choose the voice from a catalog — some voices available even on the free plan, more as the plan goes up.
A 7-Point Checklist for Choosing a Read-Aloud Tool
- Japanese naturalness. Listen to a sample as long as your own text; judge by ear, not spec sheets
- Character limits. How much text per generation — PUBVOICE takes up to 5,000 characters per run
- Voice variety. How many voices, and how many on the free tier
- Emotional expression. Whether emotion/tone control exists, and which plan includes it
- Commercial use. Whether generated audio may be used in published or distributed work — check the terms
- Data handling. Whether your submitted text is kept out of AI training
- Pricing structure. The free tier's size, and whether billing is by characters or minutes
Rank these for your own use case rather than hunting for a tool that maxes all seven. For audio files, 2 and 7 are non-negotiable; for quality, 1 and 4. And in the end, nothing beats feeding your own text and listening.
Turning Text into an MP3 with PUBVOICE
- Paste the text — up to 5,000 characters per generation; split longer drafts at section breaks
- Pick a voice from the catalog — some voices are usable on the free plan, more as the plan rises
- Insert emotion tags such as
[excited]or[whispers]where you want the tone to shift — a paid-plan feature (Lite and above), not available on the free plan - Generate and listen; fix readings of proper nouns and numbers, then regenerate
- Download as MP3 and carry it on your phone
No recording, no editing — revise the text and regenerate. PUBVOICE also has AI transcription for meetings and recordings, covered in the voice-memo transcription guide.
FAQ
Q. How far does the free tier go?
A. Built-in read-aloud features are free everywhere. PUBVOICE's free plan includes 10,000 characters of voice generation and 60 minutes of AI transcription per month, with no credit card. Emotion tags and the wider voice catalog come with paid plans.
Q. What are emotion tags?
A. Tags embedded in the text to control the reading's emotion and tone. PUBVOICE offers six — [excited], [laughs], [whispers], [sad], [curious], [shouting] — placed anywhere and combinable. Available from the Lite plan and up.
Q. Is there a character limit per generation?
A. Up to 5,000 characters per run in PUBVOICE. For longer drafts, split at section breaks and play the exported MP3s in order.
Q. Can I use generated audio commercially?
A. Commercial-use terms are defined per service in their terms of service. Check the relevant clause before using audio in published or distributed work.
Related Articles
- The Complete Guide to Turning Blog Articles into Audio — narrating your articles with AI and embedding a player on your site
- Transcribing Voice Memos and IC Recorder Audio: The Complete Guide — the reverse direction: turning recordings into text
- How to Transcribe Audio for Free — phone features, self-hosted models, and free tiers compared
- AI Transcription Services Compared (2026 Edition) — pricing, free tiers, and export formats side by side
If the text you want to hear is already on your screen, all that remains is pasting it. PUBVOICE's free plan includes 10,000 characters of voice generation and 60 minutes of transcription per month with no credit card required — take one text and turn it into audio at PUBVOICE.

Yutaro Sasao
CEO / MediaLeap Inc.
After leading web media monetization and data analytics at KADOKAWA / DWANGO, and driving programmatic ad revenue growth in SSP / ad network businesses, he founded MediaLeap Inc. in May 2025. He now develops and operates AI audio SaaS "PUBVOICE", tourism DX app "ANIME TRAVEL", and AI voice chat app "AITOMO". Drawing on cross-functional expertise in advertising, technology, analytics, and business, he works to improve media revenue through data-driven strategies.
Related posts
Transcribing Voice Memos and IC Recorder Audio: The Complete Guide from File Export to AI Transcription
A complete guide to transcribing audio from iPhone Voice Memos and IC recorders: recording tips, exporting m4a and mp3 files, and choosing among built-in phone features, self-hosted Whisper, and AI transcription services.
AI Transcription Services Compared (2026): A Practical Guide for Individuals and Teams
Notta, toruno, Rimo Voice, Otolio, AI GIJIROKU, YOMEL, and PUBVOICE compared on free tiers, effective per-hour pricing, diarization, editing, and export formats — with task-by-task recommendations and annual cost examples.
Meeting Minutes, How to Write Them Faster: A Copy-Paste Template and an AI Transcription Workflow
A practical guide to writing meeting minutes: a copy-paste template that puts decisions and to-dos up top, plus the full workflow of recording, AI transcription, editing, and sharing, including the honest limits of AI.