PUBVOICE.
    FEATURESPRICINGFAQHELP
    Log inStart Free
    P
    PUBVOICE.

    AUDIO_CONTENT
    PLATFORM V1.0

    ◆ ALL SYSTEMS OPERATIONAL

    PRODUCT

    • What is PUBVOICE
    • Blog

    COMPANY

    • About Us

    LEGAL

    • Terms of Service
    • Privacy Policy
    • Legal Notice

    SUPPORT

    • Help Center
    • FAQ
    • Contact Us
    STATUS

    ALL SYSTEMS ONLINE

    Ready!

    SERVICES

    💬

    AITOMO

    AI voice chat and image generation with your favorite characters

    🌏

    Kaigai Matome

    Overseas reactions to Japan, curated and translated by AI

    🍜

    RAMEN TRIP

    Find, seal, and master every bowl of ramen

    📷

    TOKYO LENS

    An AI guide that explains Tokyo through your camera

    © 2026 PUBVOICE. All rights reserved.

    Made with care in Japan

    A Field Comparison of Audio Embed Services for Web Media

    A Field Comparison of Audio Embed Services for Web Media

    July 29, 2026|Updated: August 16, 2026

    Table of Contents

    1. 1.Voice quality alone won't decide the winner
    2. 2.The "audio CMS" pattern that BeyondWords established
    3. 3.Sorting services into three tiers makes the choice visible
    4. ·Category A: Article-audio platform type
    5. ·Category B: Infrastructure / accessibility type
    6. ·Category C: Technical component type
    7. 7.Comparison matrix
    8. 8.If your goal is to lift retention rates
    9. 9.If your goal is to restore ad revenue
    10. 10.The white space left in the Japanese market

    More media teams are hitting a ceiling on ad revenue and starting to ask, "Can we turn our articles into audio?" The number of services to choose from has grown, too. ElevenLabs, ReadSpeaker, CoeFont, BeyondWords, and Japan's PUBVOICE. Lined up by name alone, they all look like "audio."

    An article-audio service for web media refers to an end-to-end system that automatically converts articles into speech, embeds a player on the site, measures playback data, and in some cases ties into audio advertising for monetization. The dimensions worth comparing for a simple TTS (text-to-speech) tool versus a publisher-grade audio CMS are completely different.

    Since this is a comparison piece, let me state the conflict of interest up front. I currently build and sell an article-audio SaaS called PUBVOICE. I have field experience on both the ad and the media sides, but the goal here is to lay out a framework for service selection; the explanation of our own product's performance numbers I will leave to a separate article.

    Voice quality alone won't decide the winner

    Before the comparison, let me set the axes first.

    Japanese TTS engines (AITalk, CoeFont, ElevenLabs, and each vendor's neural voices) have already reached a practical level. If you are not used to them, there is room for taste preferences. However, the phase where micro-differences in voice quality alone determine "which service a media company chooses" is probably behind us.

    ReadSpeaker offers more than 90 languages and over 300 neural voices, and ElevenLabs excels at multilingual support and voice cloning. Technical voice quality has reached a practical range across all the major services.

    What actually matters in the field is how easy it is to deploy, scripting that understands article structure, playback measurement, and whether you can connect it all to audio advertising. If we place BeyondWords as the benchmark, the comparison axes boil down to these four.

    Four comparison axes for choosing an article-audio service: ease of deployment, scripting, measurement, and monetization

    The "audio CMS" pattern that BeyondWords established

    BeyondWords is an audio CMS widely used by overseas publishers. It has a track record with major media such as News Corp, and it is common to hear that it delivers audio articles to several million people per week. According to Vendr (a crowdsourced aggregation of procurement data, so the sample is limited), around $3,000 per year is the average contract price, with enterprise pricing going higher.

    The benchmark (the ideal feature set) in this article is the following five points:

    • Audio generation: High-quality AI voices plus voice clones of reporters and editors
    • Workflow integration: Automatic article detection and audio generation from WordPress or RSS
    • Player embedding: Place an audio player on the website via a JS tag, etc.
    • Analytics: Play-start rate, completion rate, listener behavior (GA4 integration, etc.)
    • Monetization: Sponsor audio insertion before and after audio articles, ad-server integration

    BeyondWords's strength is that monetization is built in as standard. "Generate an audio article → let listeners play it in the player → earn with audio ads" all lives inside a single product. What Japanese media probably want to imitate is this end-to-end pattern.

    Sorting services into three tiers makes the choice visible

    Services roughly fall into three tiers, based on how comprehensively their features serve a media company's goal (retention improvement or ad monetization).

    Category A: Article-audio platform type

    A tier that covers automatic article audio generation, player embedding, and playback analytics within a single product. Whether monetization is built in varies by product.

    BeyondWords (overseas)
    For large publishers. The monetization module is standard, making it the closest to the BeyondWords-style "end-to-end" model. Around $3,000 per year on average. Whether operation and support in a Japanese-language environment work well should be verified before deployment.

    PUBVOICE (Japan)
    Can be deployed with RSS registration and a single line of JavaScript embedding. It understands article headings and bullet lists, rewrites them into an ear-friendly script, and then generates audio. GA4 integration visualizes play-start and near-completion behaviors. A free plan is available, and paid plans are published from 1,480 JPY/month to 29,800 JPY/month. Currently centered on article audio and measurement; audio ads are not built in (see below).

    BotTalk (overseas)
    An audio distribution platform for publishers. Marketed as being able to turn web content into audio without API integration. One of the services aimed at overseas publishers.

    Category B: Infrastructure / accessibility type

    These skew toward "making the site listenable" as infrastructure. They have many deployments. Audio advertising and monetization workflows for media tend to require a separate build by design.

    ReadSpeaker webReader (overseas / Japan)
    Over 12,000 deployments worldwide (per the company; timing to be confirmed). A single tag audio-enables all articles. Offers over 90 languages and more than 300 neural voices. Japanese-language stability and accessibility compliance (Act on Elimination of Discrimination against Persons with Disabilities, JIS X 8341-3:2016, etc.) are strengths. A strong option for media that prioritize full-article read-aloud and accessibility.

    Web Yomi Shokunin (Japan)
    A product offered by AI Inc. Audio-enables via HTML tag embedding. Widely adopted by local governments and for accessibility use. Monetization and analytics requirements for media typically need to be checked case by case within the scope of the deployment plan and surrounding integrations.

    Category C: Technical component type

    A tier of services whose features are specialized in one area, serving as parts for building a full stack in-house.

    ElevenLabs / Audio Native (overseas)
    Strong on voice quality, multilingual support, and cloning. Audio Native also supports article embedding. Analytics and monetization must be built in-house. An option as a high-quality TTS engine.

    Amazon Polly (overseas)
    AWS-based. The story that The Washington Post used it in a custom implementation is well known. Strictly for full in-house builds.

    CoeFont (Japan)
    Over 10,000 AI voices and an API. A Japanese voice technology layer. Players and ad integration must be developed in-house.

    Otonaru Inc. (Japan)
    Based on public information, its main field is audio advertising, podcast production, and ad-server integration. It tends to make the shortlist when the topic in Japan is "building audio ad inventory and sales." Its product design philosophy differs from article-embedding audio SaaS. Whether it includes TTS or automatic article audio may vary by product and plan, so confirmation against your requirements is necessary.

    Comparison matrix

    I summarized the axes that matter when selecting in the field into a table.

    Service Ease of deployment Retention analytics Ad monetization Suitable use case (rough guide) Notes
    BeyondWords High (CMS integration) Standard feature Standard feature Overseas publishers / end-to-end Around $3,000/year, benchmark
    PUBVOICE High (RSS + 1 line of JS) GA4 integration Currently measurement-focused Japanese media / article audio Free plan available; paid from 1,480 JPY/month
    BotTalk High To confirm To confirm Overseas publishers / auto audio No public information
    ReadSpeaker High (tag) Depends on product Custom design is common Accessibility / full-article read-aloud 90 languages, 300 voices; large deployment base
    Web Yomi Shokunin High (tag) To confirm Custom design is common Local government / accessibility Provided by AI Inc.
    ElevenLabs Medium (development) In-house build In-house build High-quality TTS / in-house Embeddable via Audio Native
    CoeFont Low (API) In-house build In-house build Japanese voice API / in-house No public information
    Otonaru Inc. To confirm To confirm A primary strength Japan / audio ads / ad server Feature scope to confirm

    A classification diagram sorting audio services into three tiers: platform type, infrastructure type, and technical component type

    If your goal is to lift retention rates

    If the goal is improving session duration and browsing depth for core users, the three essential requirements are these:

    Playback data measurement, automatic article audio generation, and near-no-code deployment.

    With ReadSpeaker or AITalk you can also create a "listen-able" state. However, to spin the retention improvement cycle, an article-audio platform that ties into your web analytics tools is the better fit. Without visibility into play-start rate and near-completion metrics, the editorial side stops at "we added audio," and it becomes hard to form hypotheses for improvement.

    In the Japanese market, for media that want automatic article audio and GA4 integration as a set, products like PUBVOICE tend to enter the shortlist. The fact that it understands article structure and scripts accordingly may contribute to improving completion rates. "Not just a read-aloud" is the direction of feedback from deployment partners.

    If you aim for the same thing overseas, BeyondWords is the first candidate. How far it covers Japanese and domestic support needs to be confirmed case by case.

    If your goal is to restore ad revenue

    If the goal is new ad inventory (pre-roll, mid-roll, integration with an existing ad server), the main characters of the comparison shift to the audio CMS and audio ad infrastructure side.

    Overseas, there are examples like BeyondWords where article audio and audio ads live within a single product. In Japan too, there are several services focused on audio ad delivery design, podcast production, and ad-server integration. On public information, Otonaru Inc. is a name that sometimes enters the shortlist as a product that positions strength in that area. Feature scope can change by plan and timing, so matching requirements against official materials is mandatory.

    Meanwhile, article-embedding audio SaaS (many products including PUBVOICE) is, for now, centered on retention and measurement. Whether the design closes the loop through audio ads within the same dashboard varies by product.

    Merely installing article audio will not replace display advertising. Monetization that includes inventory design, sales, and measurement is a story you build with separate products or in-house development. If you have an in-house team, there is also a configuration where you generate audio with ElevenLabs or CoeFont and load it onto your existing ad infrastructure. The Washington Post's custom implementation with Polly is a story for a major with engineering resources; the cost for a mid-size media company to take the same path is high.

    The white space left in the Japanese market

    Voice quality alone no longer creates much differentiation. What media companies are really looking for is how much of the workflow they can outsource.

    In Japan, my impression is that options for the BeyondWords-style "end-to-end, from article audio through player to audio ads" are still limited. There are many cases where article-audio SaaS and audio-ad / ad-server-oriented services exist as separate products. So in selection, I think a realistic order is to first clarify the goal (retention or new ad inventory), and then shortlist products in the category that fits it.

    About the author and PUBVOICE

    I wrote about the motives and background in Why we built an AI audio SaaS for web media.

    PUBVOICE's current product is centered on article audio and playback measurement. It is not designed to close the loop on audio ads within the same product the way BeyondWords does, nor do we assume a bundled deployment package with other companies' services.

    On audio advertising, we are at the stage of considering delivery integration via VAST tags, among other things, so that it sits easily on top of the ad operations the media company already has. Timing and specs are undecided, and for deployment decisions at this point I think the framing of "retention and measurement as the main purpose" is appropriate.

    Leaving only the comparison axes

    Choosing a service is not a matter of voice-quality preference; it is a question of how much of the workflow you outsource.

    If you want to run a measurement-and-improvement cycle for retention, go with an article-audio platform that ties into GA4. If creating ad inventory comes first, shortlist audio-CMS- or audio-ad-infrastructure-oriented products while matching requirements against the official feature scope. If accessibility and full-article read-aloud come first, ReadSpeaker and Web Yomi Shokunin are also strong candidates.

    I wrote about the structural, ongoing decline of text display CPM in Why Web Media Ad Revenue Stagnates: Dissecting 3 Structural Problems. Audio is not a replacement for that. A large part of it is about increasing the channels through which readers can be reached.

    What media companies choose next. The right answer is probably different for each one. I will leave the comparison axes here.

    Pricing and feature information is based on publicly available information as of August 2026. Please confirm each vendor's latest official materials.

    Related articles

    • What changes when web media articles are turned into audio — Deployment data on dwell time (limited measurement)
    • Why we built an AI audio SaaS for web media — Why we did not choose an ad model and the conflict-of-interest background
    • Why Web Media Ad Revenue Stagnates: Dissecting 3 Structural Problems — The structural limits of display advertising
    • How We're Thinking About Raising the Value of Existing Ad Slots Without Adding More — A concept on dwell time and slot value
    • Imitation in the AI era and delivery beyond text — Why non-text delivery still matters

    Related posts

    +
    3x avgDwell time
    +5%Return visitors
    30+Voice patterns
    ◆ FREE_PLAN_AVAILABLE

    Try audio publishing for free.
    Experience it with no commitment.

    Try the audio experience for free with your real articles.
    All essential features are available on the Free plan at no cost.

    Get started for freeSee more features

    No credit card required · Cancel anytime

  1. 11.About the author and PUBVOICE
  2. 12.Leaving only the comparison axes
  3. 13.Related articles
  4. Yutaro Sasao

    Yutaro Sasao

    CEO / MediaLeap Inc.

    After leading web media monetization and data analytics at KADOKAWA / DWANGO, and driving programmatic ad revenue growth in SSP / ad network businesses, he founded MediaLeap Inc. in May 2025. He now develops and operates AI audio SaaS "PUBVOICE", tourism DX app "ANIME TRAVEL", and AI voice chat app "AITOMO". Drawing on cross-functional expertise in advertising, technology, analytics, and business, he works to improve media revenue through data-driven strategies.

    ← Back to blog

    How We're Thinking About Raising the Value of Existing Ad Slots Without Adding More

    Can we raise the value of existing ads without adding slots? A former 'add more slots' practitioner writes a working concept on dwell time, viewability, and attention metrics.

    Imitated Content in the AI Era—and Delivery Beyond Text

    With AI in 74% of new pages, text imitation costs near zero. What remains for creators is thought, context, and non-text delivery—voice, pacing, presence.

    What Changes When Web Media Adds Audio to Articles

    At some deployments that added AI audio to web media articles, dwell time averaged 3x, repeat rate +5%, and PV +5%. A practitioner from the media monetization field writes about the market data and the barriers to adoption.