Guides

How to listen to arXiv papers (and what any text-to-speech tool gets wrong)

arXiv is close to the best case for listening to research. Nearly every paper on it was built from LaTeX source, so the PDF has real, selectable text rather than a scanned image, and a read-aloud extension can start on it immediately — no upload, no conversion. It's also the worst case in one specific way: equations, citation markers and figure captions read badly in every text-to-speech tool ever made, including ours. This guide covers both halves: the setup, the abstract-first workflow, the HTML versions that often listen better than the PDF, and an honest list of what to expect.

Listen to a paper in three steps

  1. Install a read-aloud extension. We make Lullula, free for Chrome, Edge and Firefox. It reads text-based PDFs and web pages, which covers both formats arXiv offers. No account needed to start.
  2. Open the paper in a tab. On the abstract page, the PDF link opens the paper in your browser's own viewer. If the abstract page also shows an HTML link, that version is usually the cleaner listen — more on that below.
  3. Press play. The paper is read aloud in a natural voice, each sentence is highlighted as it's spoken so you can follow a dense argument with your eyes, and your place is saved. Close the tab at the end of the method section and resume there tomorrow.

Speed is adjustable from 0.5× to 2×. Most people settle at 1.25× to 1.5× within a week; past that you're skimming with your ears — fine for related work, wrong for the part you need.

Add Lullula to Chrome — Free Works in Chrome and Edge · Firefox version available

The abstract-first workflow: summary, then decide

A reading list of thirty papers is not thirty listens. An 8,000-word paper is about 55 minutes at normal speed, so the expensive mistake isn't listening slowly, it's listening to the wrong paper. The workflow that works: read the abstract on the arXiv page first — it's short, and it's the authors' own summary. If it's promising, open the full paper and ask for an AI summary; thirty seconds tells you whether the method section matters for your question or whether the abstract already gave you everything. Only the survivors get the full listen.

One data note: generating a summary sends the paper's text to Lullula's servers, as does playing a premium voice. With your browser's built-in voices, nothing leaves your machine, and the PDF itself is never uploaded either way.

PDF or HTML? The HTML version is often the cleaner listen

Since late 2023, arXiv has generated HTML versions of new papers from their LaTeX source. When one exists, the abstract page shows an HTML link next to the PDF link, and the address is arxiv.org/html/ followed by the paper's ID. For listening it's usually the better choice: a single column instead of two, no page headers or page numbers creeping into the audio, and figure captions in their own blocks rather than interrupting a paragraph mid-sentence. It reads like any other web page.

The caveats: not every paper has one — the conversion fails on some LaTeX, and older papers were never converted — so the PDF stays the fallback. And HTML is no kinder to equations, which brings us to the honest part.

Where every text-to-speech tool struggles: equations, citations, captions

Equations. Typeset math is symbols positioned on a page, not a sentence. A voice will read fragments of it — a letter here, "equals" there — or skip it. Nobody has solved this, and we won't pretend we have. What helps is the highlighting: when the voice hits an equation, you can see which one, glance at it yourself, and let playback move on to the prose that explains it.

Citation markers. "[12]" and "(Vaswani et al., 2017)" are read literally, and a related-work section with forty of them is noisy. If you know the field, that's the section to skip.

Figure captions and reading order. In a two-column PDF, a caption can land in the middle of a paragraph, and the reading order across columns can occasionally surprise you. The HTML version avoids most of this, which is the main reason to prefer it.

Tables and the reference list. A table read aloud is a sequence of numbers with no rows; look at it instead. And when the voice reaches the references, stop. Nobody wants forty of them narrated.

Long papers and surveys: resume across sessions

A 40-page survey is a week of listening, not an afternoon. Your place is saved per document, so you can close the tab at the end of a section and pick up from the same sentence the next day, and you can bookmark the papers you're working through to a reading list so the queue lives in one place. For book-length PDFs, the PDF guide covers the page-boundary quirks.

Ways to listen to research papers, compared

OptionWhat you getBest when
Lullula (that's us) Reads the PDF or HTML version in place, with sentence highlighting, saved place and summaries. Unlimited listening in browser voices plus 2,000 premium words a month, free, no account. Papers live in your browser tabs and you want to read along with the highlight.
Paper-to-audio apps You upload the PDF, their servers process it, and you get a podcast-style feed on your phone. Some try to handle references and equations more gracefully than a plain reader. You want a phone-first queue and don't mind uploading, which matters for unpublished or licensed papers.
Edge's built-in Read Aloud Free, no install, and it reads PDFs in Edge's viewer. No saved place, no summary, whole-page reading. You use Edge and listen occasionally.
Your operating system Narrator (Windows) or Spoken Content (macOS) read anything on screen, free. No install at all, and you can live with a system voice and no highlighting.

Frequently asked questions

Does arXiv have an audio version of papers?

No. arXiv publishes the PDF, the source files and, for many recent papers, an HTML version. Turning any of those into audio is up to you, which is what the steps above are for.

Can it read the equations?

Not usefully, and neither can any other text-to-speech tool. Expect fragments or silence at each equation, use the highlight to see which one it's on, and read the math with your eyes. The prose around it, which is most of a paper, reads fine.

Should I listen to the PDF or the HTML version?

HTML when it exists: single column, no page headers, captions kept out of the paragraphs. PDF when it doesn't, which is the case for older papers and for some newer ones whose LaTeX didn't convert.

How long does a paper take to listen to?

Figure 130 to 150 spoken words a minute at normal speed. An 8,000-word paper is about 55 minutes at 1×, or under 40 minutes at 1.5×.

Is the paper uploaded anywhere?

The file, no: it opens in your browser. When you play a premium voice or request a summary, the text of that page is sent to Lullula's servers to produce the audio or the summary. With the browser's built-in voices, nothing leaves your machine. We don't track your browsing or sell data; the privacy policy has the specifics.

What about papers from journals, or old scanned articles?

Journal PDFs with selectable text work the same way: drag the file into a tab and press play. A scanned article that's only images of pages won't read, because there's no OCR built in. Run it through an OCR tool first to add a text layer.

Start with the paper at the top of your list

Open its abstract page, click the HTML link if there is one, and press play on the introduction. Within a section you'll know whether hearing papers suits you, and the free tier is enough to find out. The full detail on PDF support is on the Read PDFs aloud page.

Add Lullula to Chrome — Free No account · No card · 2,000 premium words free every month