By iwantfreepdftools.com Team

Studying From PDFs: What the Research Says

By iwantfreepdftools.com Team · 28 min read

Quick Answer

PDF tools help students study by turning scattered readings into one organised, searchable set of notes. The useful ones do three jobs: merge and split files so each subject lives in one place, compress files to fit upload limits, and pull text out for active recall practice. Free browser tools handle all three.

PDF tools help students study by doing three unglamorous jobs really well: merging scattered readings into one file per subject, splitting huge textbooks down to the chapter you were actually assigned, and pulling the text out so you can turn it into questions you test yourself on. That last one is where the real gain is. The file management saves you minutes. The active recall saves you the exam. And the good news is that everything here is free, and most of it runs in the browser tab you already have open.

Advertisement

What do PDF tools actually do for students?

Most of your course material arrives as a PDF now, and that shift is not slowing down. In a survey of 3,447 faculty and administrators fielded in April 2025 by Bay View Analytics, close to 30% of faculty said they require digital-only textbooks compared with just 10% who require print only, and 77% agreed that digital materials give students more flexibility. That reporting comes via Inside Higher Ed. Translation: you are going to spend a lot of your degree inside PDF files, so it is worth being good at it.

PDF tools split into two families, and students need both. The first is readers and annotators, the apps where you actually read, highlight, and scribble in the margin. The second is file tools, the utilities that reshape the document itself: merge, split, compress, convert, extract. Readers get all the attention. File tools are the ones that quietly stop you losing an afternoon to a submission portal that will not accept your 40MB scan.

Here is the thing nobody tells first-years. The single biggest time sink is not reading. It is hunting. You have eleven PDFs open across three folders, two downloads called "week4_final_FINAL.pdf", and no idea which one had the diagram you needed. Fixing that is a five minute job once, and it pays off every week after.

Does studying from PDFs work as well as paper?

Slightly less well, and it is worth knowing why so you can compensate. Pablo Delgado and colleagues at the University of Valencia ran a meta-analysis published in Educational Research Review in 2018 that pooled 54 studies and more than 170,000 participants. As summarised by the Technical University of Munich Center for Educational Technologies, reading on paper beat reading on screen with an effect size of g = -0.21.

Two details matter more than the headline. First, the paper advantage showed up for expository texts, the explanatory kind you get in textbooks and journal articles, but not for narrative texts. Second, the gap widened when reading time was limited. And it did not shrink over the years studied, so the idea that digital natives grow out of it does not hold up.

So what do you do with that? Not print everything. The practical fix is to remove the two things screens make worse: time pressure and shallow skimming. Give yourself longer than you think a reading needs, and force yourself to do something with the text rather than sliding your eyes down it. But before the highlighter, there is a question that meta-analysis pools right over: which screen?

Does it matter which device you read your PDFs on?

Yes, but less than you'd guess, and the group it hits hardest is not the one you'd expect.

The Delgado meta-analysis above pooled screens of every kind. Six years later the same research group narrowed the question to the devices students actually carry. Ladislao Salmerón, Lidia Altamura, Pablo Delgado, Anastasia Karagiorgi and Cristina Vargas published a review and meta-analysis of reading comprehension on handheld devices versus paper in the Journal of Educational Psychology in 2024, indexed by ERIC as EJ1409878.

Paper still won, and the margin was small. Across 38 between-participant comparisons the effect was g = -0.113, and across 21 within-participant comparisons it was g = -0.103. Put that next to the g = -0.21 from the broader 2018 analysis and the handheld gap is roughly half the size. Tablets and phones aren't uniquely bad for reading. They're screens, and screens cost you a little.

Then comes the finding that should stop you. Undergraduates showed stronger screen inferiority than primary and secondary students. Not weaker.

That's the opposite of the story everyone tells about digital natives growing into their devices. And it lines up with what the 2018 analysis already found, that the gap held steady across the years studied rather than fading. Before you take it too personally, there's a likely explanation sitting in the earlier result: the paper advantage shows up for expository text and not narrative text. University reading is almost entirely expository. So part of what looks like an age effect is probably a difficulty effect, and you're reading exactly the kind of text that screens handle worst.

One more moderator is worth your attention. Screen inferiority was larger when students were assessed individually than in groups. Which is to say it shows up most in the situation you're actually revising for.

What a phone does differently

Phones have been studied on their own, and the mechanism is stranger than "small screen, harder to read". Motoyasu Honma and colleagues ran a crossover experiment with 34 Japanese university students, published in Scientific Reports in 2022. Everyone read on both media. Comprehension scores were lower on the smartphone, with a main effect of medium at F(1,32) = 26.445 and p < 0.0001.

The interesting part is what else they measured. Students sighed less while reading on the phone, and activity in the prefrontal cortex went up compared with paper. Their path analysis tied the two together and on to performance, linking left-channel activity to the number of sighs (p = 0.021) and to reading scores directly (p = 0.003). The authors read this as sigh inhibition and prefrontal overactivity combining to drag comprehension down. You're working harder and breathing shallower, and you don't notice either.

Treat that as a lead rather than settled fact. It's 34 people in one lab reading novel excerpts, which is a long way from your reading list. But it points the same direction as the meta-analysis, and it explains something you've probably felt without naming it.

So what should you actually do?

Keep the size of all this in perspective, though. An effect of g = -0.11 is small, and it's dwarfed by the difference between reading passively and testing yourself. Nobody ever failed because they used a tablet. They failed because they highlighted for four hours and never closed the file. Which brings us to the highlighter.

Why does zooming in on your phone make a PDF worse, not better?

Because a PDF is a fixed layout by design, and zooming does not change that. You get bigger text and a page that no longer fits, so you end up panning left and right on every single line.

This is the mechanism behind the device findings above. The meta-analyses tell you handheld screens cost you a little comprehension. They do not tell you why, and part of the answer is not about screens at all. It is about what a PDF is.

A web page reflows. Make the text bigger and the words rewrap into the narrower column. A PDF holds its page geometry no matter what, because it was built to print the same everywhere. That is the whole point of the format, and on a 6-inch screen it becomes the problem.

There is a useful benchmark for how bad that is. The W3C's Success Criterion 1.4.10 Reflow asks that content be presentable "without loss of information or functionality, and without requiring scrolling in two dimensions" for vertically scrolling content at a width equivalent to 320 CSS pixels. It carves out an exception for things that genuinely need a two-dimensional layout, like maps and diagrams. It is a web standard rather than a rule about your lecture handout, but as a test of whether a document is readable on a small screen, it is the right question. Most academic PDFs fail it outright.

Quick test: open the PDF on your phone and zoom until the body text is comfortable. If you now have to swipe sideways to finish a line, that file is going to be miserable for a two-hour reading. Read it on something bigger, or convert it.

And the two-column academic paper is the worst case of all. Zoom enough to read one column and the other one is off-screen, so you scroll down the left column, back up to the top, then down the right. On a phone that is dozens of round trips per page, and every one of them is a chance to lose your place.

Now the part that matters more than convenience. If you need large text to read at all, this stops being an annoyance and becomes a barrier. The student who most needs to enlarge the type is the one for whom fixed layout does the most damage, which is the same population the locked-file section above is about.

What actually helps, in rough order of effort.

None of this means phone reading is a mistake. It means the format is fighting you on small screens in a specific, predictable way, and knowing the shape of it lets you pick your battles rather than assuming you are just bad at concentrating.

Does it matter what else is open while you read?

More than the device does. And the size of the effect is bigger than most of the paper-versus-screen gap this guide has already covered.

Yamin Shen pooled the evidence in Distractions in digital reading: a meta-analysis of attentional interference effects, published in Frontiers in Psychology in 2025. It covers 32 empirical studies and 124 independent samples running from 2000 to 2025, comparing comprehension when something else is competing for attention against comprehension when nothing is.

The overall result: Hedges' g of −0.64, with a 95 percent confidence interval of −0.89 to −0.40 (t = −5.16, p < 0.001). Negative means worse. That's a medium-to-large hit to how much of the reading you actually take in.

Three things in the moderator analysis are worth more than the headline number.

The device wasn't the problem. Reading device came out as a non-significant moderator, alongside text type, time constraints and publication year. So the distraction penalty doesn't care whether you're on a laptop, a tablet or a phone. This is a genuinely useful separation from the section above: your phone is a worse place to read a PDF because of screen size and reflow, not because phones are inherently distracting surfaces. A laptop with six tabs open is the same problem in a bigger window.

Music was as bad as television. Both landed at g = −0.82, the joint worst category in the analysis. Interruptions, by contrast, came in at just g = −0.21, the mildest thing tested.

That ordering is the opposite of what most people assume. A flatmate asking you something costs you relatively little, because you deal with it and return. Continuous background audio never lets you return, because it never stops competing. If you have a study playlist running while you work through a course reading, that's the finding to sit with.

The moderator that should worry you: the penalty got worse with age, not better. College students dropped by g = −0.72. School-age readers, from elementary through high school, came in between roughly −0.20 and −0.27. Whatever tolerance you think you've built up for reading with things running in the background, the data points the other way.

One honest caveat on the numbers. Between-group studies produced a much larger effect (g = −0.81) than within-group designs (g = −0.32), and that design difference was itself statistically significant at p = 0.015. So the true size depends a fair bit on how you measure it. The direction is not in doubt. The precise magnitude is softer than a single headline figure suggests.

There's an older finding that adds something the meta-analysis can't, and it matters if you study anywhere near other people.

Faria Sana, Tina Weston and Nicholas Cepeda ran a simulated classroom study published as Laptop multitasking hinders classroom learning for both users and nearby peers in Computers & Education in 2013, indexed by the US Department of Education's ERIC database. They report that participants who multitasked on a laptop during a lecture scored lower than those who didn't, and that participants who were in direct view of a multitasking peer scored lower than those who were not.

So the cost isn't only yours. Sitting where someone else's screen is in your field of view does measurable damage to what you take in, and you never chose to open those tabs. If you read in a library or a shared study space, where you sit is a real variable.

What to actually do with this, keeping it proportionate:

And notice how this reframes the paper-versus-screen material earlier in this guide. Part of the reason paper tends to win in those studies is that a sheet of paper cannot show you a notification. Strip the distractions out of the digital version and you close some of that gap without buying a printer.

Why does highlighting alone not help you remember?

Because it feels like learning without being learning. John Dunlosky and colleagues reviewed ten of the most common study techniques for Psychological Science in the Public Interest in 2013, and as Kent State University reported, only two earned a high utility rating: practice testing and distributed practice. Highlighting and underlining landed in the low utility group, next to rereading and summarization. Those are, of course, exactly the techniques students use most.

The problem is that dragging a yellow bar across a sentence takes almost no thought. You end up with a page that looks studied and a brain that has not retrieved anything. Retrieval is the bit that builds memory, and highlighting skips it entirely.

There's a study that explains why the wrong method feels so convincing. Henry Roediger and Jeffrey Karpicke had students read passages, then either reread them or take recall tests, and checked retention at three points. Their results, published as Test-Enhanced Learning in Psychological Science in 2006, split by timing. After 5 minutes, rereading won. After 2 days and after a week, testing won by a wide margin.

Now the part that should change how you study. Rereading also increased students' confidence that they would remember the material, while leaving them worse off a week later. So the technique that feels most productive in the moment is the one quietly failing you, and your own sense of how well a session went is not a reliable guide. If you have ever finished a highlighting session feeling on top of a topic and then blanked in the exam, that gap is the whole finding.

It also explains why cramming the night before appears to work. You're testing yourself at the 5 minute end of that curve, where rereading genuinely does look better. The bill arrives later.

The fix: highlight while you read, then close the PDF and write down what the highlighted section said from memory. Check yourself afterwards. That one change converts a low utility technique into practice testing, which is the highest rated technique in the review.

This is where PDF text extraction earns its place in a study workflow. Pull the text out of a chapter, strip it to the key claims, and convert each one into a question. Our PDF to text tool does the extraction in your browser, and the full text extraction guide covers what to do when the PDF is a scan rather than real text.

Is highlighting ever worth doing?

Yes, but not for the reason you'd assume, and not in the way most students do it. Nine years after the Dunlosky review, Héctor Ponce, Richard Mayer and Ester Méndez went back and pooled the evidence properly. Their meta-analysis in Educational Psychology Review in 2022, indexed by the US Department of Education's ERIC database, covered 36 published articles and 85 effect sizes. It splits the question in a way a blanket low utility verdict can't.

When students highlighted for themselves, it improved memory with an average effect size of 0.36. Comprehension barely moved, at 0.20.

Read that as two different jobs. Remembering what a chapter said is not the same as understanding the argument it makes. Highlighting does something for the first and close to nothing for the second. So if your exam is definitions, dates and formulas, your highlighter is pulling some weight. If the question starts with "critically evaluate", it isn't.

Then there's the age split, which almost nobody quotes. Learner-generated highlighting worked for college students at 0.39 and much less for school students at 0.24. If you picked up the habit at fifteen and someone has since told you it's worthless, that's why the research looks contradictory. At university you're better at spotting which sentence carries the load, and the technique improves along with you.

But the finding that should actually change how you handle your files is the other half of the paper. Instructor-provided highlighting, where somebody else marked the text before you ever read it, averaged 0.44 for memory and comprehension alike. That beats doing it yourself on both counts. And it held up across the board: 0.41 for college students, 0.48 for school students.

Someone else's highlights outperform your own because they carry a judgement you haven't made yet. Your lecturer already knows which paragraph the whole argument rests on. You're still working that out while you drag the cursor, which is exactly when your highlighting is least informed.

What this means for your PDFs: a reading your lecturer already annotated, a chapter with the key passages marked, or a classmate's marked-up copy is worth more than a clean file you highlight from scratch. Merge those into your topic file instead of starting over with a fresh download. Then still do the recall step on top, because none of these numbers come close to the effect of testing yourself.

Ready to sort out your course files?

Open Free PDF Tools →

All in your browser. No signup, no upload.

Should you type your notes or write them by hand?

You've probably seen the headline: handwriting beats typing, science says so. It's worth looking at what that study actually did, because the answer is more useful than the headline.

Van der Weel and Van der Meer at the Norwegian University of Science and Technology put 36 university students in a 256-channel EEG cap and had them write words with a digital pen and type the same words on a keyboard. Published in Frontiers in Psychology, they found far richer brain connectivity during handwriting: enhanced theta (3.5 to 7.5 Hz) and alpha (8 to 12.5 Hz) coherence across parietal and central regions, with 16 significant connections for handwriting against essentially none for typing.

That sounds decisive. But another set of researchers read the same paper and disagreed in print, which is the part almost nobody quotes.

Pinet and Longcamp published a commentary in the same journal in 2025 arguing the evidence doesn't support the learning claim. Their objections are specific:

So the honest state of play is that handwriting produced a different brain signature in a lab, and whether that translates into remembering your reading better is unproven.

What to take from it: don't buy a paper notebook because of an EEG study. The variable that keeps showing up across all of this research is effortful re-expression. Writing a claim in your own words is what does the work, and it does the work whether your hands are on a pen or a keyboard. Copying a sentence out is weak either way.

Which lands in the same place as the highlighting section. The method matters less than whether you had to reconstruct the idea to produce it. If typing tempts you to transcribe the slide verbatim, that's a real argument for slowing down, and handwriting is one way to force it. Just know that's the mechanism, not magic in the pen.

Should you listen to your readings instead of reading them?

For getting through volume, yes, and it holds up better than most people expect. For the readings you'll have to argue about, no. The research splits along a line you've already met earlier on this page.

Every PDF reader worth using will read a document aloud, and most phones will too. If you have a commute, a gym session or 400 pages of reading you're never going to finish sitting down, that's a real option rather than a cop-out.

Virginia Clinton-Lisell pooled the evidence in Review of Educational Research (2022), covering 46 studies and 4,687 participants. The headline is permission-giving: overall, reading and listening comprehension were not reliably different (g = 0.07, p = .23).

So as a way of getting content into your head, listening is not the inferior option it's assumed to be. That alone should change how you handle a reading list you're behind on.

But the moderators are where the actual advice lives, and there are two.

Self-pacing matters. Reading came out ahead when readers set their own pace (g = 0.13, p = .049). That makes sense. When you read, you slow down at the hard paragraph and go back over the sentence you didn't get. Audio just keeps going, and you have to actively decide to rewind.

The type of understanding matters more. For inferential comprehension, meaning working out what follows from the text rather than what it stated outright, reading won clearly (g = 0.36, p = .02). For literal comprehension, catching what the text actually said, there was nothing in it at all (g = -0.01, p = .93).

Read that against the highlighting section above and it's the same line drawn twice. Highlighting helped memory at 0.36 and did close to nothing for comprehension at 0.20. Listening matches reading on literal recall and falls behind on inference. Two completely different research literatures, landing on the same distinction: getting the content in is one job, and reasoning about it is another.

The sorting rule: listen to the readings you need to have covered. Read the ones you'll be asked to evaluate, compare or critique. If the essay question starts with "critically", that PDF gets your eyes.

There's a separate question hiding inside this one, and it deserves its own answer. If reading itself is hard for you, does read-aloud help? That isn't the same as asking whether listening matches reading for a fluent reader.

Wood, Moxley, Tighe and Wagner ran a meta-analysis on exactly that for the Journal of Learning Disabilities (2018), looking at text-to-speech and related read-aloud tools for students with reading disabilities. They found an average weighted effect size of 0.35, with a 95 percent confidence interval of 0.14 to 0.56 (p < 0.01). Small to moderate, and reliably above zero.

The authors note that study design explained some of the variance, which is the usual caution about a mixed literature. But the direction is clear enough: if decoding text is the bottleneck, having it read to you removes a barrier rather than lowering the bar.

Two practical notes before you set this up.

It only works on real text. A scanned lecture PDF is a picture of a page, so there's nothing for the reader to say out loud. That's the same wall you hit when you try to search or highlight one, and the fix is the same, covered in the scanned PDFs section below.

And the recall step still applies. This is the trap. Listening feels like effort because it takes an hour, but the effort is passive, exactly like dragging a highlighter. Nothing on this page changes the central finding: input isn't what builds the memory, retrieval is. So finish a listening session the same way you'd finish a reading one. Stop, and say out loud what the chapter argued before you check.

Should you let AI summarise your PDFs?

For getting through a reading list, it's tempting and it works. For actually learning the material, there's now a controlled trial showing exactly how it backfires, and the mechanism is the same one this whole article keeps circling.

Bastani and colleagues ran the study everyone in education has been waiting for, and its title gives away the finding: "Generative AI without guardrails can harm learning: Evidence from high school mathematics," in the Proceedings of the National Academy of Sciences (2025). Hold onto those two words, without guardrails, because the study is not a verdict on AI. It's a verdict on unsupervised AI. Nearly 1,000 high school maths students were split across three conditions during practice sessions: no AI, a standard ChatGPT-style interface they call GPT Base, and GPT Tutor, the same model wrapped in prompts that gave hints instead of answers.

During practice, AI looked spectacular. GPT Base users scored 48 percent higher than the control group. GPT Tutor users scored 127 percent higher.

Then the researchers took the AI away and ran an exam.

The GPT Base group scored 17 percent worse than students who never had AI at all. Not worse than the tutor group. Worse than having nothing. Practising with an assistant that hands over answers left them less able than if they'd struggled alone. In the GPT Tutor condition, where the model withheld answers and offered hints, that damage was largely wiped out.

Read the two numbers together: a 48 percent gain during practice and a 17 percent loss on the exam came from the same tool in the same students. If you judge a study method by how the session felt, unguarded AI is the most convincing thing you will ever use. That is precisely the trap the rereading research described earlier, with a faster engine.

The authors' word for it is crutch. And it maps exactly onto everything above. Retrieval builds memory. Effortful re-expression builds understanding. An AI that reads the chapter and hands you five tidy bullet points has performed the retrieval and done the re-expressing, and you have done neither.

Two caveats before you delete your account. This was high school mathematics, where the answer is either right or wrong, and it may not transfer cleanly to reading a dense article in your own field. And the tutor condition genuinely worked, which is the more useful half of the finding: the problem isn't the model, it's what you ask it to do.

So use it deliberately rather than not at all:

If you want the tool-by-tool picture rather than the learning science, our sister site SpotFreeAI covers free AI tools for students.

Which free PDF tools should students actually use?

You need a reader, a set of file tools, and a citation manager. That is genuinely the whole stack.

JobFree optionWhy it fits student work
Read and annotateFoxit, PDFgear, Xodo, Microsoft EdgeHighlight, sticky notes, and drawing without a subscription
Merge readingsMerge PDFOne file per week or per topic instead of a download folder
Pull one chapterSplit PDFGrab pages 240 to 268 and leave the other 600 behind
Fit upload limitsCompress PDFGets scanned work under portal caps
Extract for recallPDF to TextTurns readings into question material
Edit a draftPDF to WordReworks a PDF handout into editable text
Manage citationsZoteroFree, stores the PDF next to the reference, generates bibliographies

Zotero deserves a specific mention because it is the one non-obvious pick. It is free, it is run by a non-profit, and it attaches the actual PDF to each reference. When your supervisor asks where a claim came from in March, you are one click from the highlighted page rather than one hour from giving up.

That table is deliberately short, because this page is about how to study from a PDF rather than how to operate the tools. Step-by-step instructions for each job, where a submission size cap actually comes from, and a semester workflow built around the tools themselves all live on the companion page, free PDF tools every student needs. Go there if you have a file to fix today. Stay here if you want to know what the evidence says about reading, highlighting and remembering.

How do you build a PDF study workflow that sticks?

Simple beats clever. A workflow you follow badly still beats a beautiful system you abandon in week three.

  1. One folder per module. Not per week, not per lecturer. Per module.
  2. Rename on download. Use W04-Topic-Author.pdf. Ten seconds now, no hunting later.
  3. Merge weekly. At the end of each week, merge that week's readings into one file. You get a single searchable document per topic.
  4. Split what is oversized. If a textbook is 700 pages and you need 30, split them out. Smaller files open faster and search faster.
  5. Highlight, then close and recall. The step everyone skips. It is the step that works.
  6. Extract to questions. Once per topic, pull the text and write ten questions. Those become your revision deck.

Spread that recall practice out rather than cramming it. Distributed practice was the other high utility technique in the Dunlosky review, and it costs nothing to apply. Ten questions revisited four times across a month beats forty questions the night before.

When should you split, merge, or compress a study PDF?

Split when the file is bigger than what you need. Textbooks, full journal issues, and past paper archives all fall here. Working from a 30 page extract instead of a 700 page book makes searching faster and stops you scrolling past your own bookmarks.

Merge when related material is scattered. Lecture slides plus the two assigned papers plus your own typed notes, combined into one file, means a single search box covers the whole topic. It also means one file to open on the train.

Compress when something has to be uploaded or emailed. Most submission portals sit between 10MB and 25MB, and phone-scanned pages blow straight past that because cameras save at far higher resolution than anyone reading the page will ever need. Our guide on compressing a PDF for email walks through the target sizes and what actually shrinks.

One warning on splitting textbooks. Copyright rules still apply to your personal copies, and sharing extracts around a group chat is a different thing from making your own reading easier. Check what your institution's licence allows.

Why can't you search or highlight some lecture PDFs?

Because it isn't really text. It's a photograph of text, and your PDF reader has no more idea what it says than it would of a holiday snap.

You've met this file. Someone scanned a chapter from a library book, or a lecturer photographed their handwritten notes, and the result opens fine but fights you. Ctrl+F finds nothing. You can't drag-select a sentence. The highlighter draws a yellow box that isn't attached to any words. Every study technique in this article assumes you can get at the text, and this one file type quietly blocks all of them.

The distinction has a name. Section508.gov, the US federal accessibility programme, calls the thing you're missing renderable or searchable text, and it's blunt about the consequence: without it, "individuals who rely on assistive technology, such as screen readers, will be unable to read or interact with the content."

Which turns a personal annoyance into something bigger the moment you share the file. According to the National Center for Education Statistics, 21 percent of undergraduates reported having a disability in 2019 to 2020. The survey's categories include blindness or serious difficulty seeing, deafness or serious difficulty hearing, serious difficulty concentrating, remembering or making decisions, and serious difficulty walking or climbing stairs, so the first three are the ones a scanned page bears on directly. That's roughly one in five people in your cohort, and 11 percent of postgraduates. Drop a scanned chapter into the group chat and for some of them you've shared a blank page.

The fix is OCR, optical character recognition, which Section508.gov describes as software with "the ability to convert scanned content into searchable text." It reads the shapes on the image and works out which letters they are. Run it once and the same file becomes searchable, selectable, highlightable and readable aloud.

Three things worth knowing before you rely on it:

Our PDF to text tool handles the extraction side once a file has real text in it, and the PDF accessibility guide covers tagging and structure if you're producing documents other people have to read rather than just consuming them.

Why does a PDF refuse to let you copy or annotate it?

Because there are two completely different reasons a PDF fights you, they look identical from the outside, and they need opposite responses.

The section above covers the first one. No text layer, because the file is a photograph of a page, and the fix is OCR. But there's a second failure that catches people out badly, and running OCR on it wastes an hour for nothing.

Some PDFs carry permission settings. The text layer is perfectly fine, sitting right there in the file, but the document has flags set that tell readers to block copying, printing, or annotating. Library ebooks do this. So do publisher course packs, exam papers and plenty of institutional handouts.

Here's the ten second test that tells you which one you've got.

Get this diagnosis right and you save yourself a lot of pointless work. OCR cannot unlock a permissions flag, and no amount of re-exporting will change a file that has no text in it to begin with.

What can you do if a locked file blocks your screen reader?

Go through your institution rather than around the lock. And the law is more helpful here than most students realise.

This stops being an inconvenience and starts being a barrier the moment you depend on assistive technology. A permissions flag that blocks text extraction can also stop a screen reader getting at the words, which means the file is not merely annoying, it's unusable.

Two separate pieces of US law address that.

The first is 17 U.S.C. 121, known as the Chafee Amendment. It says it is not an infringement of copyright for an authorized entity to reproduce or distribute copies of a previously published literary work in accessible formats, exclusively for eligible persons. The statute defines an eligible person as someone who is blind, or who has a visual impairment or reading disability that prevents reading printed works to substantially the same degree as a person without that disability, or who is otherwise unable through physical disability to hold or manipulate a book or to focus or move the eyes. An authorized entity means a nonprofit organisation or a governmental agency whose primary mission is providing specialised services relating to training, education, or the adaptive reading or information access needs of people with disabilities. Congress updated that language in 2018 through the Marrakesh Treaty Implementation Act, which is also where section 121A and its cross-border provisions came from.

The second is the anti-circumvention side. Section 1201 of the copyright act normally makes it unlawful to break a technological measure that controls access to a work, but the Librarian of Congress runs a rulemaking every three years that carves out exceptions. In the ninth of those proceedings, in a final rule effective 28 October 2024, one of the adopted classes covers literary works or previously published musical works fixed as text or notation, distributed electronically, that are protected by technological measures which either prevent the enabling of read-aloud functionality or interfere with screen readers or other assistive technologies.

Two caveats matter before you rely on any of that.

Which points at what you should actually do, and it isn't downloading a cracking tool.

And if this is a file you're producing rather than consuming, don't set permissions you don't actually need. Our PDF accessibility guide covers what to do instead.

How do you quote and cite accurately from a PDF?

Carefully, because two things go wrong quietly. The text you copy out isn't always the text on the page, and the page number your reader shows you often isn't the page number you should cite. Both produce errors that look fine in your draft and fall apart if anyone checks.

The section above dealt with scans that have no real text at all. This is the subtler problem: files that do have text, where the extraction is still imperfect.

Why copied text doesn't always match the page

A PDF stores the position of individual characters, not words or paragraphs. Reassembling those into sentences is guesswork, and the guessing gets harder as layouts get more complex.

Adhikari and Agarwal measured this properly, comparing ten parsing tools across six document categories in arXiv preprint 2410.09871 (2024). On financial documents the best performer hit an F1 score of 0.9885, essentially clean. On scientific articles the same tool managed 0.8526, and the authors note all tools showed a marked decrease in that category.

Which is a problem, because scientific articles are exactly what most of your reading list is made of. The formats that break extraction hardest are the ones you're citing.

The failure modes they document are worth recognising on sight:

None of these announce themselves. A quote with a silently deleted ligature or a rejoined hyphen is a misquotation, and if a marker compares it against the source, that's what it looks like.

So the rule is simple. Never paste a quotation straight from a PDF into an essay without reading it against the page. Copy, paste, then look at the original and fix what moved.

Which page number should you actually cite?

The one printed on the page, not the one in your viewer's toolbar. Those numbers agree less often than you'd hope, because front matter, cover sheets and repository banners all shift the count.

The Bodleian Libraries at Oxford make the case for why PDFs are the format worth having: ebooks in PDF "are the best to use as almost all retain the original layout and pagination of the print copy." That's the whole advantage. It only pays off if you read the number off the page itself.

Their guidance for the awkward case is useful too. Where no page numbers exist, use paragraph numbering instead, counting paragraphs from the start of the chapter. And they flag that EPUB and Kindle location numbers are not a substitute, because those shift with font size, line spacing and margins.

One more trap specific to academic PDFs. A preprint or accepted manuscript is usually paginated differently from the published version. If you cite page 12 of a preprint and your reader opens the journal version, page 12 is a different passage. Check which version you actually have before you quote a page from it.

What to record while you still have the file open

Do this at the point of reading and your reference list assembles itself. Leave it to the week before submission and you'll be reopening forty files to find one page number.

Can you study from PDFs without uploading your files?

Yes, and for coursework you should. A lot of free PDF sites work by uploading your document to their server, processing it there, then deleting it after some retention window you have to take on trust. For a public journal article that is fine. For your unmarked dissertation draft, a scan of your student ID, or a paper under embargo, it is a decision you should make on purpose rather than by accident.

Browser-based tools avoid the question entirely. Every tool on this site runs inside your own browser using JavaScript, so the file never leaves your laptop. There is no upload, no account, and nothing sitting on a server waiting to be deleted. That also means they keep working on campus wifi that is having a bad day, which is more useful than it sounds during deadline week.

And when you finish the degree and start applying for things, our sister site IWantFreeResume.com handles the graduate CV side for free too.

What else do students ask about PDF study tools?

What is the best free PDF tool for students?

There is no single best one, because students need three different jobs done. For reading and annotating, use a free reader like Foxit, PDFgear, Xodo, or the viewer built into Microsoft Edge. For file surgery like merging, splitting, and compressing, use browser tools that run on your own device. For citations, Zotero is free and stores your PDFs alongside your references.

Is highlighting a PDF a good way to study?

Not on its own. Dunlosky and colleagues reviewed ten common study techniques in Psychological Science in the Public Interest in 2013 and rated highlighting and underlining as low utility, alongside rereading and summarization. Highlighting only helps when it feeds something else, such as turning your highlights into practice questions you test yourself on later.

How do I make a PDF small enough to upload to my LMS?

Compress it first, and if that is not enough, split it. Most submission portals cap uploads somewhere between 10MB and 25MB. Image-heavy files shrink the most because the images are usually saved at far higher resolution than a screen or a printer needs. Scanned pages photographed on a phone are the usual culprit.

Can I study from PDFs on my phone?

Yes, but treat the phone as a review device rather than a first-read device. Small screens force more scrolling, which makes it harder to hold the structure of an argument in your head. Read new material on a laptop or tablet, then use your phone for going back over highlights and notes between classes.

Are free online PDF tools safe for coursework?

It depends on whether the tool uploads your file. Many free sites send your document to a server, process it there, and keep it for a period before deleting it. Browser-based tools that run entirely on your own device never transmit the file at all, which matters if you are handling unpublished research, graded work, or anything with your student ID on it.

Sources: Adhikari, N. S. and Agarwal, S. (2024), "A Comparative Study of PDF Parsing Tools Across Diverse Document Categories," arXiv:2410.09871, ten parsing tools across six document categories, for the F1 of 0.9885 on financial documents against 0.8526 on scientific articles, the marked decrease across all tools in the scientific category, and the documented failure modes of broken words, mishandled hyphenation, mangled special characters and disrupted word sequence in multi-column layouts. Bodleian Libraries, University of Oxford, Citing ebooks guide, for PDFs retaining the original layout and pagination of the print copy, the recommendation to use paragraph numbering counted from the start of the chapter where no page numbers exist, and the unreliability of EPUB and Kindle location numbers because they shift with font size, line spacing and margins. DOI Foundation, "What is a DOI?", for the DOI as a persistent, long-lasting reference and for metadata being updatable without changing the identifier. Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). "Improving Students' Learning With Effective Learning Techniques," Psychological Science in the Public Interest, via Kent State University (kent.edu). Ponce, H. R., Mayer, R. E., & Méndez, E. E. (2022). "Effects of Learner-Generated Highlighting and Instructor-Provided Highlighting on Learning from Text: A Meta-Analysis," Educational Psychology Review, 34(2), 989-1024, indexed by the US Department of Education's ERIC database (EJ1334669), for the 36 articles and 85 effect sizes, the 0.36 memory and 0.20 comprehension figures for learner-generated highlighting, the 0.39 college and 0.24 school split, and the 0.44 instructor-provided average with its 0.41 and 0.48 breakdown. Roediger, H. L. III, & Karpicke, J. D. (2006). "Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention," Psychological Science, 17(3), 249-255 (sagepub.com), for the finding that repeated studying beat repeated testing at a 5 minute retention interval while testing produced substantially greater retention at 2 days and 1 week, and for the result that repeated studying nonetheless increased students' confidence in their ability to remember the material. Delgado, P., Vargas, C., Ackerman, R., & Salmerón, L. (2018). "Don't throw away your printed books," Educational Research Review, via TUM Center for Educational Technologies (edtech.tum.de). Salmerón, L., Altamura, L., Delgado, P., Karagiorgi, A. and Vargas, C. (2024). "Reading Comprehension on Handheld Devices versus on Paper: A Narrative Review and Meta-Analysis of the Medium Effect and Its Moderators," Journal of Educational Psychology, ERIC EJ1409878 (eric.ed.gov), for g = -0.113 across 38 between-participant comparisons, g = -0.103 across 21 within-participant comparisons, and the moderators showing stronger screen inferiority among undergraduates than school students and under individual rather than group assessment. Honma, M. et al. (2022). "Reading on a smartphone affects sigh generation, brain activity, and comprehension," Scientific Reports (PMC8803971), for the 34 participants, the main effect of medium at F(1,32) = 26.445 with p < 0.0001, and the sigh and prefrontal cortex path analysis. Bay View Analytics faculty survey, April 2025, reported by Inside Higher Ed (insidehighered.com). Van der Weel, F. R. and Van der Meer, A. L. H. (2024). "Handwriting but not typewriting leads to widespread brain connectivity: a high-density EEG study with implications for the classroom," Frontiers in Psychology (frontiersin.org), for the 36 participants and the theta and alpha connectivity findings. Pinet, S. and Longcamp, M. (2025). Commentary on the same paper, Frontiers in Psychology (PMC11750765), for the methodological objections. Both are cited so you can weigh them yourself. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö. and Mariman, R. (2025). "Generative AI without guardrails can harm learning: Evidence from high school mathematics," Proceedings of the National Academy of Sciences (PubMed 40560616), for the roughly 1,000 students, the 48 and 127 percent practice gains, and the 17 percent exam drop in the unguarded condition. A correction was issued for that paper in 2025, but it fixed an author affiliation only and left the findings unchanged. US General Services Administration, Section508.gov, "Converting Scanned Documents into Section 508 Conformant PDFs", for the description of scanned pages as problematic for assistive technology, the definition of renderable or searchable text, and OCR as the conversion step. US Department of Education, National Center for Education Statistics, Fast Facts: Students with Disabilities, for the 21 percent of undergraduates and 11 percent of postbaccalaureate students reporting a disability in 2019 to 2020, and the survey's definition of disability. Clinton-Lisell, V. (2022), "Listening Ears or Reading Eyes: A Meta-Analysis of Reading and Listening Comprehension Comparisons," Review of Educational Research, indexed by the US Department of Education's ERIC database as EJ1347325, covering 46 studies and 4,687 participants, for the overall comparison (g = 0.07, p = .23), the self-paced reading advantage (g = 0.13, p = .049) and the split between inferential comprehension (g = 0.36, p = .02) and literal comprehension (g = -0.01, p = .93). Wood, S.G., Moxley, J.H., Tighe, E.L. and Wagner, R.K. (2018), "Does Use of Text-to-Speech and Related Read-Aloud Tools Improve Reading Comprehension for Students With Reading Disabilities? A Meta-Analysis," Journal of Learning Disabilities, ERIC EJ1164251, for the average weighted effect size of 0.35 with a 95 percent confidence interval of 0.14 to 0.56 (p < 0.01) and the note that study design moderated some of the variance. Shen, Y. (2025), "Distractions in digital reading: a meta-analysis of attentional interference effects," Frontiers in Psychology 16, DOI 10.3389/fpsyg.2025.1671214, for the 32 empirical studies and 124 independent samples covering 2000 to 2025, the overall Hedges' g of −0.64 with a 95 percent confidence interval of −0.89 to −0.40 (t = −5.16, p < 0.001), the g of −0.82 for both television and music against −0.21 for interruptions, the age split of −0.72 for college students against roughly −0.20 to −0.27 for elementary through high school readers, the non-significant moderator result for reading device, and the significant research-design difference (between-group g = −0.81, within-group g = −0.32, p = 0.015) that makes the precise magnitude less settled than the direction. Sana, F., Weston, T. and Cepeda, N.J. (2013), "Laptop multitasking hinders classroom learning for both users and nearby peers," Computers & Education 62, 24 to 31, ERIC EJ1007624, for the findings that laptop multitaskers scored lower than non-multitaskers and that participants in direct view of a multitasking peer scored lower than those who were not. That study reports direction rather than effect sizes in its abstract, and it used a simulated classroom rather than a live course.

Get fresh reads straight to your inbox

Get notified when we publish new articles. Unsubscribe anytime.