A Field Comparison of Audio Embed Services for Web Media
More media teams are hitting a ceiling on ad revenue and starting to ask, "Can we turn our articles into audio?" The number of services to choose from has grown, too. ElevenLabs, ReadSpeaker, CoeFont, BeyondWords, and Japan's PUBVOICE. Lined up by name alone, they all look like "audio."
An article-audio service for web media refers to an end-to-end system that automatically converts articles into speech, embeds a player on the site, measures playback data, and in some cases ties into audio advertising for monetization. The dimensions worth comparing for a simple TTS (text-to-speech) tool versus a publisher-grade audio CMS are completely different.
Since this is a comparison piece, let me state the conflict of interest up front. I currently build and sell an article-audio SaaS called PUBVOICE. I have field experience on both the ad and the media sides, but the goal here is to lay out a framework for service selection; the explanation of our own product's performance numbers I will leave to a separate article.
Voice quality alone won't decide the winner
Before the comparison, let me set the axes first.
Japanese TTS engines (AITalk, CoeFont, ElevenLabs, and each vendor's neural voices) have already reached a practical level. If you are not used to them, there is room for taste preferences. However, the phase where micro-differences in voice quality alone determine "which service a media company chooses" is probably behind us.
ReadSpeaker offers more than 90 languages and over 300 neural voices, and ElevenLabs excels at multilingual support and voice cloning. Technical voice quality has reached a practical range across all the major services.
What actually matters in the field is how easy it is to deploy, scripting that understands article structure, playback measurement, and whether you can connect it all to audio advertising. If we place BeyondWords as the benchmark, the comparison axes boil down to these four.

The "audio CMS" pattern that BeyondWords established
BeyondWords is an audio CMS widely used by overseas publishers. It has a track record with major media such as News Corp, and it is common to hear that it delivers audio articles to several million people per week. According to Vendr (a crowdsourced aggregation of procurement data, so the sample is limited), around $3,000 per year is the average contract price, with enterprise pricing going higher.
The benchmark (the ideal feature set) in this article is the following five points:
- Audio generation: High-quality AI voices plus voice clones of reporters and editors
- Workflow integration: Automatic article detection and audio generation from WordPress or RSS
- Player embedding: Place an audio player on the website via a JS tag, etc.
- Analytics: Play-start rate, completion rate, listener behavior (GA4 integration, etc.)
- Monetization: Sponsor audio insertion before and after audio articles, ad-server integration
BeyondWords's strength is that monetization is built in as standard. "Generate an audio article → let listeners play it in the player → earn with audio ads" all lives inside a single product. What Japanese media probably want to imitate is this end-to-end pattern.
Sorting services into three tiers makes the choice visible
Services roughly fall into three tiers, based on how comprehensively their features serve a media company's goal (retention improvement or ad monetization).
Category A: Article-audio platform type
A tier that covers automatic article audio generation, player embedding, and playback analytics within a single product. Whether monetization is built in varies by product.
BeyondWords (overseas)
For large publishers. The monetization module is standard, making it the closest to the BeyondWords-style "end-to-end" model. Around $3,000 per year on average. Whether operation and support in a Japanese-language environment work well should be verified before deployment.
PUBVOICE (Japan)
Can be deployed with RSS registration and a single line of JavaScript embedding. It understands article headings and bullet lists, rewrites them into an ear-friendly script, and then generates audio. GA4 integration visualizes play-start and near-completion behaviors. A free plan is available, and paid plans are published from 1,480 JPY/month to 29,800 JPY/month. Currently centered on article audio and measurement; audio ads are not built in (see below).
BotTalk (overseas)
An audio distribution platform for publishers. Marketed as being able to turn web content into audio without API integration. One of the services aimed at overseas publishers.
Category B: Infrastructure / accessibility type
These skew toward "making the site listenable" as infrastructure. They have many deployments. Audio advertising and monetization workflows for media tend to require a separate build by design.
ReadSpeaker webReader (overseas / Japan)
Over 12,000 deployments worldwide (per the company; timing to be confirmed). A single tag audio-enables all articles. Offers over 90 languages and more than 300 neural voices. Japanese-language stability and accessibility compliance (Act on Elimination of Discrimination against Persons with Disabilities, JIS X 8341-3:2016, etc.) are strengths. A strong option for media that prioritize full-article read-aloud and accessibility.
Web Yomi Shokunin (Japan)
A product offered by AI Inc. Audio-enables via HTML tag embedding. Widely adopted by local governments and for accessibility use. Monetization and analytics requirements for media typically need to be checked case by case within the scope of the deployment plan and surrounding integrations.
Category C: Technical component type
A tier of services whose features are specialized in one area, serving as parts for building a full stack in-house.
ElevenLabs / Audio Native (overseas)
Strong on voice quality, multilingual support, and cloning. Audio Native also supports article embedding. Analytics and monetization must be built in-house. An option as a high-quality TTS engine.
Amazon Polly (overseas)
AWS-based. The story that The Washington Post used it in a custom implementation is well known. Strictly for full in-house builds.
CoeFont (Japan)
Over 10,000 AI voices and an API. A Japanese voice technology layer. Players and ad integration must be developed in-house.
Otonaru Inc. (Japan)
Based on public information, its main field is audio advertising, podcast production, and ad-server integration. It tends to make the shortlist when the topic in Japan is "building audio ad inventory and sales." Its product design philosophy differs from article-embedding audio SaaS. Whether it includes TTS or automatic article audio may vary by product and plan, so confirmation against your requirements is necessary.
Comparison matrix
I summarized the axes that matter when selecting in the field into a table.
| Service | Ease of deployment | Retention analytics | Ad monetization | Suitable use case (rough guide) | Notes |
|---|---|---|---|---|---|
| BeyondWords | High (CMS integration) | Standard feature | Standard feature | Overseas publishers / end-to-end | Around $3,000/year, benchmark |
| PUBVOICE | High (RSS + 1 line of JS) | GA4 integration | Currently measurement-focused | Japanese media / article audio | Free plan available; paid from 1,480 JPY/month |
| BotTalk | High | To confirm | To confirm | Overseas publishers / auto audio | No public information |
| ReadSpeaker | High (tag) | Depends on product | Custom design is common | Accessibility / full-article read-aloud | 90 languages, 300 voices; large deployment base |
| Web Yomi Shokunin | High (tag) | To confirm | Custom design is common | Local government / accessibility | Provided by AI Inc. |
| ElevenLabs | Medium (development) | In-house build | In-house build | High-quality TTS / in-house | Embeddable via Audio Native |
| CoeFont | Low (API) | In-house build | In-house build | Japanese voice API / in-house | No public information |
| Otonaru Inc. | To confirm | To confirm | A primary strength | Japan / audio ads / ad server | Feature scope to confirm |

If your goal is to lift retention rates
If the goal is improving session duration and browsing depth for core users, the three essential requirements are these:
Playback data measurement, automatic article audio generation, and near-no-code deployment.
With ReadSpeaker or AITalk you can also create a "listen-able" state. However, to spin the retention improvement cycle, an article-audio platform that ties into your web analytics tools is the better fit. Without visibility into play-start rate and near-completion metrics, the editorial side stops at "we added audio," and it becomes hard to form hypotheses for improvement.
In the Japanese market, for media that want automatic article audio and GA4 integration as a set, products like PUBVOICE tend to enter the shortlist. The fact that it understands article structure and scripts accordingly may contribute to improving completion rates. "Not just a read-aloud" is the direction of feedback from deployment partners.
If you aim for the same thing overseas, BeyondWords is the first candidate. How far it covers Japanese and domestic support needs to be confirmed case by case.
If your goal is to restore ad revenue
If the goal is new ad inventory (pre-roll, mid-roll, integration with an existing ad server), the main characters of the comparison shift to the audio CMS and audio ad infrastructure side.
Overseas, there are examples like BeyondWords where article audio and audio ads live within a single product. In Japan too, there are several services focused on audio ad delivery design, podcast production, and ad-server integration. On public information, Otonaru Inc. is a name that sometimes enters the shortlist as a product that positions strength in that area. Feature scope can change by plan and timing, so matching requirements against official materials is mandatory.
Meanwhile, article-embedding audio SaaS (many products including PUBVOICE) is, for now, centered on retention and measurement. Whether the design closes the loop through audio ads within the same dashboard varies by product.
Merely installing article audio will not replace display advertising. Monetization that includes inventory design, sales, and measurement is a story you build with separate products or in-house development. If you have an in-house team, there is also a configuration where you generate audio with ElevenLabs or CoeFont and load it onto your existing ad infrastructure. The Washington Post's custom implementation with Polly is a story for a major with engineering resources; the cost for a mid-size media company to take the same path is high.
The white space left in the Japanese market
Voice quality alone no longer creates much differentiation. What media companies are really looking for is how much of the workflow they can outsource.
In Japan, my impression is that options for the BeyondWords-style "end-to-end, from article audio through player to audio ads" are still limited. There are many cases where article-audio SaaS and audio-ad / ad-server-oriented services exist as separate products. So in selection, I think a realistic order is to first clarify the goal (retention or new ad inventory), and then shortlist products in the category that fits it.
About the author and PUBVOICE
I wrote about the motives and background in Why we built an AI audio SaaS for web media.
PUBVOICE's current product is centered on article audio and playback measurement. It is not designed to close the loop on audio ads within the same product the way BeyondWords does, nor do we assume a bundled deployment package with other companies' services.
On audio advertising, we are at the stage of considering delivery integration via VAST tags, among other things, so that it sits easily on top of the ad operations the media company already has. Timing and specs are undecided, and for deployment decisions at this point I think the framing of "retention and measurement as the main purpose" is appropriate.
Leaving only the comparison axes
Choosing a service is not a matter of voice-quality preference; it is a question of how much of the workflow you outsource.
If you want to run a measurement-and-improvement cycle for retention, go with an article-audio platform that ties into GA4. If creating ad inventory comes first, shortlist audio-CMS- or audio-ad-infrastructure-oriented products while matching requirements against the official feature scope. If accessibility and full-article read-aloud come first, ReadSpeaker and Web Yomi Shokunin are also strong candidates.
I wrote about the structural, ongoing decline of text display CPM in Why Web Media Ad Revenue Stagnates: Dissecting 3 Structural Problems. Audio is not a replacement for that. A large part of it is about increasing the channels through which readers can be reached.
What media companies choose next. The right answer is probably different for each one. I will leave the comparison axes here.
Pricing and feature information is based on publicly available information as of August 2026. Please confirm each vendor's latest official materials.
Related articles
- What changes when web media articles are turned into audio — Deployment data on dwell time (limited measurement)
- Why we built an AI audio SaaS for web media — Why we did not choose an ad model and the conflict-of-interest background
- Why Web Media Ad Revenue Stagnates: Dissecting 3 Structural Problems — The structural limits of display advertising
- How We're Thinking About Raising the Value of Existing Ad Slots Without Adding More — A concept on dwell time and slot value
- Imitation in the AI era and delivery beyond text — Why non-text delivery still matters
