DEV Community

Lora Henley
Lora Henley

Posted on Originally published at naturaltts.org

What Students Actually Mean When They Ask for an Audio Version

The request arrives compressed. A student emails the accessibility office, or drops a line in a course forum, or tells a lecturer after class: is there an audio version of this? It reads like a yes/no question about a file that either exists or doesn't.

It isn't. "Audio version" is a category label that five or six different populations use for five or six different needs. The people fielding these requests learn this quickly, usually by producing one audio file, sending it out, and discovering that a third of the recipients stop using it within a week. The file was fine. It just wasn't the thing most of them were asking for.

Breaking the request apart is worth doing carefully, because the shape of the demand determines whether any response to it works.

The students with documented accommodations

This is the group that most institutional processes are built around, and the smallest of the populations described here. A student has documentation of a print disability — blindness, low vision, a diagnosed reading disability, a motor impairment that makes page handling difficult — and a formal accommodation attached to their record. The request routes through a defined channel and someone is accountable for fulfilling it.

What this group needs is often the least like a polished audiobook. Many of these students are experienced assistive technology users. They already have a screen reader configured to their preferences, running at a speed that sounds unintelligible to anyone else. What they need is not audio; it's accessible text — a document whose structure is intact, whose headings are real headings, whose images have descriptions, whose reading order matches the visual order. Give them a clean file and their own tools do the rest.

Where they need actual audio is at the edges: a scanned PDF that never had a text layer, a diagram-heavy chapter, a course pack that arrived as photographs of pages. The failure mode here is a document that technically exists in digital form but resists extraction. That's a preparation problem, and it lands on the same desk as the audio requests, which is part of why the two get conflated.

The students with no documentation and no diagnosis

Considerably larger, and mostly invisible to the systems designed to serve the first group. These are students who read slowly, lose the thread on long paragraphs, or find that a page of dense prose takes three passes before it means anything. Some have an undiagnosed reading difficulty. Some have attentional patterns that make sustained silent reading expensive. Some are exhausted.

They do not have an accommodation letter and frequently will not seek one. Diagnosis costs money and time, carries stigma in some contexts, and requires a student to first suspect there is something to diagnose. What they do instead is ask, informally, whether there's an audio version — often framing it as a convenience rather than a need, because that framing is safer.

This group has a distinct requirement: they typically want audio and text together. Reading along while listening is the pattern that works, because the audio sets a pace and the text anchors it. A standalone MP3 with no accompanying document is less useful to them than to almost anyone else. They also tend to want short segments. A ninety-minute single file for a full chapter is a commitment they won't make; the same chapter split at section boundaries is something they'll actually work through.

Because they don't route through formal channels, the volume of this demand is systematically undercounted. It shows up as informal asks to individual instructors, which never reach the accessibility office and never appear in any tally.

The people with no free hands

Commuters. Parents. Students working night shifts between classes. People in clinical placements, apprenticeships, or field programmes where the reading has to happen somewhere other than a desk.

Their constraint is time and physical situation, not reading ability. They read fine; they just don't have thirty uninterrupted minutes in a chair. Audio converts otherwise dead time — a bus route, a commute, a shift break — into something usable.

What this group needs is close to a podcast. Download-and-go, playable on a phone, resilient to poor connectivity, and comfortable at elevated speed. They want longer continuous files, not shorter ones, because stopping to pick the next segment while driving is a non-starter. Standard MP3 matters here more than anywhere else: whatever tool produced the audio, the output has to land in the media player they already use.

They are also the group most poorly served by a browser-based reader that streams and does nothing else. A play button on a course page assumes the student is at the course page.

The multilingual learners

A student reading academic English as a second, third, or fourth language faces a specific problem: decoding effort and comprehension effort compete for the same attention. The words are individually parseable, but the cost of parsing them leaves less capacity for the argument they're carrying.

Audio helps in several distinct ways, and multilingual learners tend to want different combinations. Some want English audio at reduced speed to hear pronunciation and prosody, using the reading as a listening exercise as much as a content delivery mechanism. Some want the material in their first language, so the concepts arrive without the language tax and the English text can be tackled afterwards with the argument already understood. Some want to alternate — first-language audio for the dense theoretical sections, English for the parts they can handle.

That means language coverage is not a nice-to-have for this population; it's the whole thing. A tool that produces only English audio serves the pronunciation use case and none of the others. Voice quality also matters more here than for native speakers, who can fill gaps from context. A flat, clipped synthetic voice removes exactly the prosodic cues that a second-language listener is relying on to parse clause boundaries.

The students who simply hear better than they read

No disability, no time constraint, no language barrier. They retain more when text arrives through their ears, and they know it about themselves.

This group is easy to dismiss as a preference rather than a need, and the learning-styles literature has taken enough hits that it's tempting to do so. But their behaviour is worth taking seriously regardless of the theory: they're the students most likely to use audio versions consistently once available, and the least likely to have any formal route to request one. They'll take whatever exists and adapt to it, which makes them a poor guide to what's actually needed — they don't complain — and a good indicator of general uptake.

Why one format satisfies almost nobody

Lay the requirements side by side and the conflicts are immediate. The documented-accommodation user wants structured text and their own voice engine. The undiagnosed reader wants short chunks synchronised with visible text. The commuter wants long downloadable files and no interface at all. The multilingual learner wants language choice and clear prosody. The auditory learner wants whatever is available.

Speed, voice, language, chunk length, and delivery format vary independently across these groups. A single MP3 of a single chapter in a single English voice hits maybe one and a half of them. This is why audio initiatives that launch with a recorded pilot — a staff member reading three chapters aloud — get warm feedback and near-zero sustained use. The recording is unimprovable by the listener. They can't speed it up past the point where the human reader's pacing collapses, can't switch language, can't get the next chapter, can't re-run it when the reading list changes.

The alternative is treating conversion as a capability rather than a deliverable: infrastructure that lets whoever prepares materials produce audio on demand, in the language and voice the requester actually needs, and hand back a plain file the student can do what they like with. That's roughly the design premise behind tools like NaturalTTS, which take pasted text or uploaded PDF and DOCX files, extract the text, and generate downloadable MP3 audio across a range of languages and voices — with a shared workspace so an accessibility team, a department, and a library aren't each solving the same request in isolation.

The demand that isn't from students at all

Two adjacent populations use the same request channel and get counted as student demand.

Researchers and postgraduates ask for audio to get through volume — papers, drafts, long-form material they need to have read rather than to have studied. Their requirements resemble the commuter's: fast, long, downloadable, disposable.

Staff ask too, and rarely in writing. Instructors and administrators who read slowly have the same needs as their students and less recourse. Learning and development teams have a different problem again: they're producing audio training content at scale rather than consuming it, and their bottleneck is throughput, not access.

Both groups get logged as "audio requests" in whatever spreadsheet exists, which further blurs the picture of who is asking for what.

The practical consequence for anyone fielding these requests is that the intake question matters more than the fulfilment. Is there an audio version is not an answerable question until you know which of these people is asking — and the person asking usually can't tell you, because they've only ever had one word for it.

Top comments (0)