PDF Compression Guide: How to Reduce File Size Free
Quick Answer
PDF compression shrinks a file's size by removing unused data and re-encoding its images. To do it free, drop your PDF into a browser-based compressor, choose a level, and download the smaller file. Text-heavy PDFs compress well with no visible change, while image-heavy ones trade some picture quality for a bigger saving.
PDF compression is the process of making a PDF file smaller by stripping out unused data and re-encoding the images inside it, and you can do it free in under a minute. The quickest way is to drop your file into a browser-based compressor, pick a compression level, and download the smaller version. Big PDFs cause real headaches, bounced emails, rejected upload forms, slow sharing, so this guide walks through what compression actually does, the difference between the lossy and lossless kinds, how to shrink a PDF for free, how much you can realistically save, and the cases where compression barely helps.
Quick option: compress a PDF in your browser right now
Compress PDF Free →No upload. Files stay on your device.
What Is PDF Compression?
Compression means encoding the same information using fewer bytes. A PDF is really a container of objects, text, fonts, images, and metadata, and the format has compression built into the standard itself. The PDF specification, standardised as ISO 32000-2 and published by the PDF Association, defines a set of filters that squeeze different kinds of content. Text and vector data get packed with a lossless filter called Flate, photos use JPEG compression, and scanned black-and-white pages use a specialised filter called JBIG2.
So a PDF is usually compressed already. When you run it through a compressor, the tool is either finding data that wasn't packed efficiently, throwing away parts nobody needs, or re-encoding the images more aggressively. The PDF format itself is documented for the long term by the United States Library of Congress, whose format description for PDF 2.0 covers ISO 32000-2:2020 and its amendments. Worth knowing if you ever want to read the actual spec: unlike most ISO standards, the PDF 2.0 bundle is available as a free download from the PDF Association rather than sitting behind the usual paywall.
Why Do PDF Files Get So Large?
Before you compress, it helps to know what's taking up the room. Three things account for almost all PDF bloat:
- Images. This is the big one. Photos and scans typically make up 80 to 90 percent of a PDF's size. A single high-resolution scanned page can be 500KB or more on its own.
- Embedded fonts. PDFs embed fonts so the file looks the same everywhere. Several font weights and styles can add hundreds of kilobytes.
- Unused objects and metadata. Files created by editing software often carry deleted objects, revision history, and extra metadata that add size without adding anything you can see.
That's why text-heavy documents like reports and contracts compress well, while image-heavy brochures and scans compress less. If you want the deeper walkthrough of specific tools, our guide on how to reduce PDF file size covers each method.
Advertisement
What's the Difference Between Lossy and Lossless Compression?
This is the core trade-off, and understanding it tells you what to expect from any tool. Both are used in PDFs, often at the same time.
| Type | How it works | Typical saving | Best for |
|---|---|---|---|
| Lossless | Removes redundancy, keeps every bit | 20 to 40 percent | Text, contracts, line art |
| Lossy | Permanently drops some image detail | 50 to 90 percent | Photos, scans, brochures |
Lossless compression tidies the data without deleting anything you'd notice, which is why text stays razor sharp. Lossy compression is more aggressive: it downsamples and re-encodes images, so you save far more but push it too hard and photos start to look soft. The smart move most compressors make is a hybrid, compress the images lossily while keeping the text and structure lossless. That gives you a big saving where the bytes actually are, without smudging your words.
There's one case where the answer is just don't. If a PDF is a record you may need to keep or hand over years from now, lossy compression can disqualify it. The US National Archives spells out the compression it will accept for digitised permanent records in its transfer guidance format tables, updated August 2025: "Deflate (ZIP), JPEG 2000 part 1 core coding system lossless compression. Agencies may use up to 20:1 visually lossless compression." That's worth reading closely, because it's more useful than a flat ban would be. Lossless is the default expectation, and the single lossy allowance comes capped and qualified: 20 to 1, and only where the loss doesn't show. You're probably not filing with the National Archives. But the principle travels: compress the copy you're emailing, and keep an untouched original for anything that counts as a record.
NARA is also useful for the one dial this whole topic actually turns on, which most guides never name: resolution. Since June 2023, its rules for digitising permanent federal records, at 36 CFR Part 1236, Subpart E, require modern textual paper records to be scanned at a minimum of 300 pixels per inch. That's the archival floor, and it's a useful reference point even though your holiday itinerary isn't a federal record.
Here's why it matters for compression. Resolution is squared, so the saving is bigger than it sounds. Halve 300 ppi to 150 and you're not storing half the pixels, you're storing a quarter of them, which is about 75 percent less image data before any re-encoding happens at all. That single change usually does more than every other setting a compressor offers combined. And 150 ppi still looks fine on a screen, because a monitor can't show you 300 ppi anyway.
So the rule of thumb: 300 ppi if the file is a record, needs printing, or might get archived. Around 150 for something people will read on screen. Below about 96 you'll start seeing it when anyone zooms in. Pick the number based on where the file is going, not on how small you wish it were.
One more thing on the lossless side, because it's about to change. The workhorse behind lossless compression in PDF has been Deflate, the same algorithm inside gzip and PNG, and it's been doing that job since the 1990s. It's fine. It's also old.
The PDF Association has published an extension adding Brotli compression to PDF as a new filter called BrotliDecode, sitting alongside the existing ones like FlateDecode for Deflate and DCTDecode for JPEG. Brotli is Google's algorithm, specified in RFC 7932 and already carrying a large share of the web's static assets. In the Association's testing it produces files averaging 20 percent smaller than Deflate, and because it's lossless, that saving costs you nothing in quality.
The format itself has moved on since that original RFC, and one of the changes lands squarely on the files you're most likely to be fighting with. RFC 9841, published in September 2025, updates RFC 7932 rather than replacing it, and adds three things: shared dictionaries, a framing format, and support for sliding windows past Brotli's original 24-bit ceiling. That last one is the interesting bit here. The RFC is explicit that the larger window brings compression gains on files above 16 MiB, which is exactly the range where a PDF stops being an inconvenience and starts getting bounced by a filing system.
Don't go looking for a Brotli button yet, though. Final inclusion in ISO 32000 is expected somewhere in 2026 or 2027, no firm date, and tool support is only starting to appear. It matters now mainly as a reason not to wreck a document chasing the last few percent. If your file is stubbornly large but the content genuinely needs to stay sharp, waiting is a real option. The same working groups are looking at JPEG XL for image data next.
What About the Fonts Inside Your PDF?
Fonts came second on that bloat list earlier and then quietly dropped out of the conversation, which is how most compression guides handle them. That's a miss, because for one category of document the fonts aren't a footnote. They're the whole problem.
PDFs embed fonts so the file looks identical on a machine that doesn't have them installed. The question is how much of the font gets embedded. Two options exist. Full embedding stores the entire typeface. Subsetting, as the Prepressure reference on PDF fonts puts it, includes "only those characters that are actually used in the layout," so if a dollar sign never appears on any page, that glyph doesn't travel with the file.
For an English document the gap between those two is real but modest. For anything in Chinese, Japanese, or Korean it's enormous, and this is the part worth knowing if you work across Asia. A CJK typeface has to cover a writing system with tens of thousands of characters. Noto Sans CJK, Google's open-source family, contains 65,535 glyphs, which happens to be the ceiling on what a single OpenType font can hold at all. Embed that whole thing to typeset an invoice using perhaps two hundred distinct characters and you've shipped the other sixty-five thousand for nothing.
That's why a two-page Chinese-language document can somehow land bigger than a fifty-page English report with photos in it. People run it through a compressor, watch the size barely move, and conclude the tool is broken. The tool is fine. It's squeezing images on a file whose weight isn't in the images.
How to tell, and what to do:
- Check the document properties. Most PDF readers list the fonts under File then Properties then Fonts. If an entry says "Embedded" rather than "Embedded Subset," the whole typeface is in there.
- Fix it upstream where you can. Word, LibreOffice, and most export dialogs have a subsetting option, often on by default but not always. Re-exporting with subsetting on is usually faster and cleaner than trying to repair the finished PDF.
- Drop fonts you don't need embedded at all. Standard faces like Arial and Times exist on virtually every device, so embedding them buys you very little. Keep embedding for anything distinctive, because that's the case the feature was built for. One exception, and it's a big one if it applies to you: don't do this to a file headed for an archive. The next section explains why.
Two honest trade-offs come with subsetting, and both are worth knowing before you turn it on everywhere. If you later need to edit the PDF and want a character that isn't in the subset, it isn't available, because it was never included. And merging two files that carry different subsets of the same font has historically produced missing or swapped characters, though modern software has largely sorted that out. Neither is a reason to avoid subsetting on a document you're sending. Both are reasons to keep a full-fidelity original if the file is one you'll edit again.
The short version: if your PDF is image-heavy, the resolution advice above is where your saving lives. If it's a text document in a CJK language and compression isn't touching it, look at the fonts before you blame the compressor.
What If the PDF Has to Be Archived or Filed Officially?
Everything above assumes the file's job is to be sent, read, and eventually forgotten. Some PDFs have a different job. They're headed into a records system, a court filing, a grant submission, or a company archive somebody will open in fifteen years. For those, the advice shifts, and one tip from the fonts section actively works against you.
Archives don't want plain PDF. They want PDF/A, a deliberately stripped-back profile of the format designed so a file renders the same decades from now without depending on anything outside itself. NARA's transfer guidance format tables list PDF/A-1 (ISO 19005-1) and PDF/A-2 (ISO 19005-2) as preferred formats for born-digital text records. Ordinary PDF, versions 1.0 through 2.0, is only acceptable. The gap between those two words is the whole idea: PDF/A takes features away from PDF on purpose, because anything clever is something that can stop working.
Fonts are the clearest case, and it's exactly where the earlier tip breaks. Standard PDF lets a file name the "base 14" fonts, Helvetica, Times, Courier, Symbol and Zapf Dingbats, without carrying them, on the assumption that every reader already has them installed. That assumption is precisely what an archivist won't sign off on. In its requirements for PDF case file collections, NARA puts it flatly: "All PDF files including both the PDF collection itself and any embedded PDF file must have all fonts, including the base 14 fonts, embedded within them."
So the saving you'd get by dropping Arial and Times is the one saving an archive specifically rules out. Worth being precise about what survives, though, because most of the font advice holds up fine:
- Subsetting still works. The requirement is that fonts be embedded, not that whole typefaces ride along. A subset is embedded. You keep the big CJK saving described above.
- Dropping the base fonts doesn't. That's the single bullet above to ignore when the destination is an archive.
- Check before you compress, not after. Plenty of compressors emit plain PDF no matter what went in, which quietly costs you conformance without warning you.
The practical workflow is to decide what a file is for before you touch a compressor. If it's a record, export to PDF/A from the source application, leave the fonts embedded, and take your size reduction from resolution instead, since that's where the bytes live anyway. If it's an email attachment, compress it however you like and keep the archival original somewhere separate. And if a file genuinely has to do both jobs, make two files. That's not a workaround. A document optimised for transmission and a document optimised for preservation are different documents, and trying to make one file serve both is how people end up with a record they can't file and an attachment that still bounces.
What If the PDF Is Going to a Commercial Printer?
Then don't compress it. Not "compress it carefully". Don't. The section above covers files that have to survive an archive, and this is its opposite number: a file that has to survive a printing press, and the rules are even less forgiving.
Here's why, and it's worth understanding rather than just taking on trust, because it explains a lot of confusing printer emails.
A print-ready PDF isn't just a PDF that happens to look good. It's a PDF built to a specification. The relevant family is PDF/X, standardised as ISO 15930, and it exists for exactly one reason: so a file can move from whoever made it to whoever prints it without a phone call. The X is for exchange.
What makes PDF/X different from an ordinary PDF is mostly a set of things it insists on. Fonts have to be embedded. Color has to be defined in a way the press can act on rather than guessed at. And the file has to declare an output intent, which is a statement inside the document saying which printing condition it was built for. That declaration is the part that matters most for this guide, because it is the thing a compressor is most likely to throw away without telling you.
So run a generic compressor over a print file and you get some combination of three problems.
- The images drop below the resolution the job needs. Compression settings are almost always tuned for screens. A press is not a screen, and a photo that looks perfect on your monitor can print visibly soft.
- The color stops being press color. CMYK and spot colors can get converted toward RGB, which is the correct choice for a screen and the wrong one for ink. Spot colors in particular do not survive a round trip.
- The output intent goes missing. And this is the binary one. Without it the file is no longer PDF/X, whatever it looks like on your screen.
That third one is why the failure is usually not gradual. It isn't that the job comes out slightly worse. It's that the printer's preflight check rejects the file before anyone looks at it, and you have lost a day.
There's a further wrinkle that explains why nobody can give you a safe universal setting. "Print" is not one requirement.
The Ghent Workgroup is the international non-profit that writes the practical rules on top of PDF/X, and it exists because, in its own account, printer and publisher associations in different countries all wanted a consistent set of rules for PDFs used to print advertisements and commercial jobs. Its current specifications are the GWG 2022 set, and they are deliberately split by market: packaging, digital print, and sign and display each get their own, because a folding carton and a magazine ad do not have the same requirements. The group also says it is preparing a future version built on PDF/X-6 and PDF 2.0.
Read that as the answer to "how much can I compress this". Your printer's spec is the answer, not a slider, and the spec depends on what is being printed.
So what do you do when the print file genuinely is too big to send?
- Ask the printer first. They deal with this daily and most of them will hand you an export preset. Thirty seconds of asking beats a rejected job.
- Fix the size at export, not afterwards. If a file is enormous because somebody placed a 300 megabyte photo and scaled it down on the page, the fix is upstream in the layout, where the resolution can be reduced to what the job actually needs rather than to what a generic compressor guesses.
- Change how you deliver it rather than what you deliver. This is the one people skip. A file too large to email is a delivery problem, not a compression problem. A shared link or the printer's own upload portal moves the whole file intact.
- If it must be split, split rather than squash. Splitting a document keeps every page exactly as it was built, which compression does not.
- Keep the compressed version for proofing only. Sending a small copy so a client can review it on a phone is completely fine. Just never let that copy be the one that reaches the press.
The ordering rule from the signatures section applies here almost word for word. Compress first and build the print file afterwards, or don't compress at all. What you cannot do is take a finished, spec-compliant deliverable and shrink it, because the compliance was the deliverable.
How Do You Compress a PDF for Free?
You don't need Photoshop or a paid app. Here's the browser route, start to finish:
- Open our free compress PDF tool. Your file stays in your browser, so nothing is sent to a server.
- Drop your PDF onto the tool or click to browse for it.
- Choose a compression level. Medium suits most files and keeps quality high.
- Let it process, then check the size reduction it reports.
- Download the smaller optimised file.
Browser-based compression is the most private option because your document never leaves your device, which matters for anything sensitive. If your goal is specifically to get a file under an email limit, our guide to compressing a PDF for email is tailored to that.
How Much Can You Reduce a PDF's Size?
It depends entirely on what's inside. A text-heavy PDF stuffed with unused metadata might drop 10 to 40 percent from structure cleanup alone. An image-heavy file can shrink 50 to 90 percent once the images are re-encoded. A clean, already-optimised PDF might barely move, and that's normal.
The number that usually matters is the email limit. According to Google's Gmail help, Gmail caps attachments at 25MB per message, and anything larger is pushed to Google Drive and sent as a link instead. Microsoft's own support pages put Outlook.com at the same 25MB, with OneDrive playing the same overflow role. So the consumer inboxes match. Where you get caught out is corporate mail: plenty of Exchange servers are still configured at 20MB or lower, and that limit is set by whoever runs the server, not by you.
There's also a hidden catch that surprises people. Email doesn't send your file as-is. It encodes attachments in Base64, which turns every 3 bytes into 4 characters, then adds a line break every 76 characters. Do the arithmetic and that's about 37 percent of inflation, so a 20MB PDF actually travels as roughly 27MB and bounces off a 25MB cap it looks like it should clear. That's the real reason aiming under 10MB is the safe target: it leaves room for the encoding and clears almost every corporate server too. If you keep bouncing, our piece on what to do when a PDF is too large to email lists five fixes.
What Size Do Official Filing Systems Actually Accept?
Whatever they say they do, and you need to look it up rather than guess. Email at least has a de facto standard at 25MB. Official submission portals have nothing of the kind, and the spread is much wider than people expect.
Two examples from the same court system make the point. The US federal courts all run electronic filing through CM/ECF, so you'd reasonably assume one limit. But the Central District of California states that "the maximum file size for PDF documents filed electronically is 35 megabytes (MB)", while the Western District of Washington puts its ceiling at 100 MB per document or attachment. Same system, same country, and a 65 MB gap between two districts.
Two things follow from that, and the second one is the useful one.
Splitting is often the sanctioned answer, not compressing harder. Washington's guidance is explicit that a document over its limit has to be divided into smaller documents or attachments, each within the limit and labelled in parts. That's the official route. Squeezing a 120 MB exhibit down to 90 MB to sneak under a cap does more damage to the evidence than filing it in two clean halves, and our guide to splitting a PDF covers doing that without disturbing anything else.
Some portals tell you the scan settings too, and they're stricter than you'd guess. The same Washington page specifies that scanned documents should be "300 dpi and black and white unless color is integral to the document", and adds that scanning at higher resolution can cause performance problems during upload.
Read that second half again, because it inverts the usual instinct. More quality is not better here. A well-meaning 600 dpi color scan of a black-and-white letter produces a file several times larger, takes longer to upload, may fail outright, and carries no extra information anybody wanted. If a portal publishes a scan spec, that spec is your compression target. You don't need to make judgement calls about downsampling at all, which is a relief given how much of this article is spent on exactly those judgement calls.
The habit worth building is boring. Find the published limit and the published format rules before you compress anything, not after your first rejection. Look for the court's or agency's own page rather than a summary of it, because these numbers change and secondary sources go stale. The California notice above is dated 2 August 2019 and is still what that district publishes.
One caution on scope. Those are two US federal district courts, picked because they publish their limits plainly. Your tax portal, immigration system, university submission tool, or tender platform will have its own numbers and its own format rules, and there's no shortcut around checking. And watch for the trap where two requirements pull against each other: a system that caps size but also demands an archival format, as covered in the section on filing above, may leave compression as the wrong lever entirely.
What Can Compression Break?
Quality is the risk everyone thinks about. It isn't the expensive one. A slightly soft photo is annoying; a document that stops working costs you a re-do. Four things are worth checking before you compress something that matters.
- Digital signatures. This one is absolute. A signature works by hashing the file at the moment of signing, so any later change to the bytes makes the hash stop matching and the signature reads as invalid. Compressing a signed PDF will break it, every time, with no way to repair it afterwards. The fix is ordering: compress first, then sign. If you've already got a signed file that's too big, you need a fresh signature on the compressed copy, not a workaround. Our digital signature guide covers how signing works.
- The invisible text layer on a scan. A scanned document that's been through OCR has two layers: the page image you see, and searchable text sitting behind it. Some compressors re-encode the images and drop that text layer, and the file looks identical while quietly becoming unsearchable. Test it the obvious way: open the compressed file and try to select a word or run a find. If nothing highlights, the layer is gone.
- Accessibility tags. A properly tagged PDF carries a structure telling a screen reader what's a heading, what's a table, and what order to read in. Aggressive optimisation can strip that structure while leaving the visible page untouched, which turns an accessible document into an unusable one for anyone using assistive tech. Our PDF accessibility guide covers what a tagged file should contain.
- Form fields. Some compressors flatten interactive forms into flat page content. The boxes still look like boxes and nobody can type in them.
None of this means don't compress. It means know what your file is before you do. A holiday-photo PDF has nothing to lose. A signed contract, a scanned archive you'll need to search, or a public document that has to meet accessibility rules all have something that can quietly stop working.
The habit that covers all four: keep the original, compress a copy, then open the copy and actually check the thing you care about. Try selecting text. Try clicking a form field. Look at whether the signature still validates. Thirty seconds of checking beats finding out when someone else does.
Can Compression Actually Change Your Content?
Everything in the section above is a thing that stops working, and a thing that stops working announces itself. This one is different, and it's the reason to be careful with scans specifically: compression can change what a document says while leaving it looking perfect.
Remember JBIG2 from earlier, the filter built for scanned black-and-white pages. Here's how it earns its very high compression ratios. It scans the page for shapes that repeat, stores one copy of each shape, and then points at that copy everywhere the shape appears again. On a page of text this is enormously efficient, because every letter "e" is very nearly the same handful of pixels. The technique has a name: Pattern Matching and Substitution.
The failure mode is sitting right there in the name. If the matcher decides two shapes are the same when they actually aren't, it substitutes one for the other, and the page it hands you looks completely clean.
This is not theoretical. In 2013 the German computer scientist David Kriesel documented it happening across a long list of Xerox WorkCentre and ColorQube machines. Sixes were coming out as eights. A cost table scanned 65 as 85. A 60 became an 80. Room dimensions on scanned building plans came out wrong. The scans looked crisp and correct, and the bug turned out to have been in the pattern-matching engine for around eight years.
Think about why that's worse than a soft photo. A blurry scan tells you it's blurry, so you know not to trust the fine detail. A substituted digit looks exactly as sharp as a correct one. There's nothing on the page to notice.
Now the fair framing, because this isn't a reason to be frightened of compression generally. That was one vendor's implementation getting the matching wrong, not proof that JBIG2 is broken, and JBIG2 also has a lossless mode that doesn't substitute anything. It's also the same instinct behind the archival caution above, where NARA sets a hard ceiling on lossy compression rather than trying to judge each file on its merits.
Three rules that follow from it:
- Numbers that matter deserve the original. Invoices, measurements, dosages, legal figures, anything where a digit changing is a real problem. Compress a copy for sending and keep the untouched scan.
- If a tool offers a "high compression" or JBIG2 mode for scans, spot-check it. Open the output next to the original and compare a page that's dense with digits. Do it once per tool, not once per file.
- Prefer dropping resolution over aggressive symbol compression. This is the useful one. Lowering resolution degrades visibly, which sounds like a downside and is actually the whole point: you can see what it cost you. Symbol substitution hides its damage, and damage you can't see is the kind that reaches someone else.
Does Compressing a PDF Remove Hidden Content?
No. Compression makes a file smaller, not cleaner, and treating one as the other is how confidential material ends up published. This is the most expensive misunderstanding on this page.
You can see why the assumption forms. A compressor rewrites the file, re-encodes the images and often drops some metadata along the way, so the thing that comes out feels freshly made. But it was built to reduce size. Removing sensitive content is a different job, and nothing in a compression pass is looking for it.
The point isn't new. The NSA published Redacting with Confidence in December 2005, on how to safely publish sanitized reports converted from Word to PDF, and its core warning was that "merely converting an MS Word document to PDF does not remove all metadata automatically." Twenty years on, the same logic covers a compression pass. Processing a file is not the same as sanitising it.
For the specifics, the US District Court for the Western District of Louisiana publishes redaction guidance for e-filers, and its list of what travels with a document is worth reading slowly: "the name and type of file, the name of the author, the location of the file on your file server, the full-sized version of a cropped picture, and prior revisions of the text."
Sit with that fourth item. The full-sized version of a cropped picture. You cropped something out, the original is still in there, and a compressor may happily re-encode it at a smaller file size while leaving it exactly as findable as it was.
The court is just as blunt about the black box, which is the mistake almost everyone makes at least once: "Highlighting text in black or using a black box over the date in MS Word or Adobe Acrobat will not protect the data from being able to be seen." The words sit underneath the rectangle. Compressing afterwards doesn't delete them. It just gives you a smaller file that still contains them.
So if a document has anything in it that shouldn't leave your desk:
- Take it out at the source. The court's own advice is to omit the sensitive material from the original document, save that version under a new name, and convert from there. Fix it before it's a PDF.
- Use a real redaction tool, not a drawing tool. Proper redaction removes the underlying objects. A shape on top of them is decoration.
- Sanitise deliberately, as its own step. Never as a hoped-for side effect of compressing, converting or flattening.
- Test it in ten seconds. Open the finished file, try to select the text under the black box, and paste it somewhere. If it pastes, you haven't redacted anything.
This belongs in a compression guide precisely because compression is the moment people expect a clean, freshly built file and don't get one. Our PDF security guide covers protecting a document properly, and digital signatures covers the related question of proving a file hasn't changed since you sent it.
What Happens to the Searchable Text in a Scanned PDF?
It can vanish, and most people only find out weeks later when a search that should work doesn't.
An OCR'd scan is really two layers stacked up. There's the picture of the page, which is where essentially all the file size lives, and behind it an invisible layer of recognised text that makes the document searchable, selectable, and readable by a screen reader. That second layer is just characters. On a long document it might be a few dozen kilobytes against many megabytes of image.
Which means there is never a size argument for throwing it away. It costs you almost nothing. But plenty of compressors flatten a PDF down to images and hand it back without the text layer anyway, and nothing in the output announces the loss. The page looks identical. It just stopped being searchable.
The order you do things in fixes most of this. OCR first, at a resolution good enough to read accurately, then compress. Going the other way round is how you get bad text: crush the image first and the recogniser is now guessing at characters you already degraded.
There's also a real distinction in how OCR writes its results, and the US National Archives has picked a side. Per NARA's transfer instructions for permanent PDF records, it "will accept PDF records that have been OCR'd using processes that do not alter the original bit-mapped image," naming Searchable Image - Exact as an example. And it "will not accept PDF records that have been OCR'd using processes that substitute OCR'd text for the original scanned text within the bit-mapped image."
That's the same worry as the section above, arriving by a different route. Pattern substitution swaps one pixel shape for another. Replacement OCR swaps the scanned image for what the recogniser thinks it said. Both leave a page that looks clean and reads wrong, which is why an archive wants the original pixels kept and the text tucked behind them rather than standing in for them.
So, three things to actually do:
- Check the output, not the input. Open the compressed file and try selecting a sentence, or search for a word you know is on page one. Takes five seconds and tells you whether the layer survived.
- Keep the image, add the text. If your OCR tool offers a choice, pick the mode that preserves the scan and adds invisible text. That's the archival-safe option and it's the safer one generally.
- Remember it isn't only about search. A scan with no text layer is a picture of words, which means a screen reader gets nothing from it at all. Our PDF accessibility guide covers what else that breaks.
One reassurance to finish, since this section has been a list of ways to lose things. Dropping image resolution, which is where your real savings come from, doesn't touch the text layer at all. The two live in different parts of the file. You can take a scan from 600 dpi to 200 and the search still works.
Does Compressing a PDF Break Its Accessibility?
It can, and the damage is invisible in the file you end up looking at. That's what makes this the most commonly missed cost of compression.
The obvious version you've already met above. Rasterise a document and every page becomes a picture, so the searchable text layer dies. Accessibility dies with it, because a screen reader hitting a rasterised page has literally nothing to read out. But there's a quieter version that catches more people. Aggressive optimise presets often include an option to discard document structure, and the structure tags are what mark your headings, lists, tables, reading order and image descriptions. None of that shows up on the page. The compressed file looks identical, opens fine, and has quietly stopped working for anyone using assistive technology.
If your documents go anywhere near a government, this isn't just a courtesy question. The US Access Board's Revised 508 Standards require, at section E205.4, that electronic content conform to Level A and Level AA of WCAG 2.0. That covers public-facing agency content and official communications, which the standards spell out as including forms, benefit notices, programme announcements and training materials. Exactly the sort of thing people compress to get under a filing limit.
There's one carve-out worth knowing, because it's narrower than people assume. E205.4 excludes non-web documents from four criteria: Bypass Blocks, Multiple Ways, Consistent Navigation and Consistent Identification. Those are navigation requirements that don't really map onto a single document. Everything else still applies, including the criteria your tags are carrying.
And here's the genuinely surprising bit, straight from the US government's own accessibility site. Section508.gov states that PDFs "are often not the most accessible or mobile-friendly option" and that federal policy recommends prioritising HTML instead. Which reframes this whole page a little. Sometimes the biggest size win on the table isn't compressing the PDF harder. It's publishing the content as a web page and leaving the PDF as an optional download. A page nobody has to download beats any compression ratio you can achieve.
When the PDF does have to ship, these are the moves:
- Check the tags before and after. Open both versions and look at the document properties or the tags panel in your reader. If the tagged flag was set and now isn't, your compression stripped it and you'd never have known from the page.
- Read the preset before you run it. Anything offering to discard document structure or document tags is offering to delete your accessibility. These options tend to be bundled into the aggressive presets and switched on by default.
- Never rasterise something people need to read with assistive tech. It's the most destructive option available, and the text layer section above covers the rest of what it costs you.
- Shrink images and subset fonts instead. Both go after the parts that actually take up the space, and neither touches the structure tree.
- If you had to go aggressive, budget time to re-tag. Adding structure back to a document afterwards takes considerably longer than compressing it did. Plan for that rather than discovering it the afternoon something is due.
When Does Compression Not Help?
Honesty helps here, because some files just won't shrink. If your PDF is a scan whose pages are already stored as compressed JPEGs, there's very little redundant data left to remove. A browser tool that only optimises structure will report a tiny or zero reduction, and that's expected, not a bug.
To actually shrink a scan, you need a tool that re-encodes the images themselves, which means either accepting some quality loss or using a heavier server-side compressor that uploads your file. Weigh that privacy trade-off against the saving. And whatever you do, keep the original, since lossy compression can't be undone. If the reason you're shrinking it is an email that keeps bouncing, our guide to compressing a PDF for email covers the provider size limits specifically.
Try our free browser-based compress tool now
Compress PDF Free →Files never leave your browser. 100% private.
Does Compressing a PDF Make It Slower to Open Online?
It can. That's the trade-off almost nobody mentions, and it's the one you can't spot by looking at the file size.
A compressed PDF that's genuinely smaller can still take longer before a reader sees anything, because two different optimisations pull in opposite directions. One makes the file small. The other makes the first page appear fast. They interfere with each other, and most compression tools quietly pick the first.
What Fast Web View actually does
The feature is called linearization, and the qpdf documentation puts the benefit in one line: with a linearized file "it is possible for a web browser to begin to display them before they are fully downloaded."
It works by restructuring the file so page one and a map of everything else sit at the front. That map is a set of hint tables. qpdf notes that it generates "only page offset, shared object, and outline hint tables", which is enough for a viewer to work out where a given page lives and request just those bytes over HTTP rather than pulling the whole document first.
This isn't an exotic modern feature either. The PDF Association's published errata for ISO 32000-2 Annex F notes that linearization "may be applied to any PDF file of version 1.2 or greater". It has been available for decades.
Where compression gets in the way
One of the better size wins in a PDF comes from object streams, which bundle many small objects together and compress them as a group. qpdf describes the contents as "compressed objects", and on a document full of structure, bookmarks, form fields, tags and annotations, that adds up.
But there's a rule. In a linearized file, per the qpdf documentation on object streams, "the linearization dictionary, document catalog, and page objects may not be contained in object streams."
Read which objects those are. The catalog and the page objects are exactly the things you'd most want bundled and compressed, and they're the ones that have to stay outside.
There's a density limit too. qpdf "never puts more than 100 objects in an object stream", which matches Acrobat's behaviour of no more than 100 objects per stream for linearized files against 200 for non-linearized ones. Linearizing halves how tightly you can pack.
Which one do you actually need?
It depends entirely on how the file reaches a reader, and the answer is usually obvious once you ask.
- Behind a download button? Ignore all of this. Linearization only pays off when a viewer streams the file. If the browser saves it and the reader opens it locally, the layout gains you nothing and the smaller file wins.
- Embedded in a page or opened in a browser viewer? Check whether your compressor preserved Fast Web View. Many strip it, because rewriting the file from scratch is how they shrink it in the first place.
- How long is it? A three-page flyer arrives before anyone notices. A 200-page report is where a reader either sees page one within a second or watches a spinner and closes the tab.
- Does your host support byte-range requests? The whole mechanism depends on a server willing to send part of a file. Without that you get the restructured layout and none of the streaming benefit.
The test that settles it takes a minute. Upload the compressed file, then open it from its URL on a normal connection, rather than double-clicking it on your desktop where everything is instant. That's the experience your readers actually get.
And keep the size of the effect in proportion. The difference between a linearized and non-linearized version of the same document is usually a few percent, not a doubling. This is not an argument against compressing. It's an argument for knowing which of the two things you needed for this particular file, because the running theme of this guide is that compression is a set of trade-offs rather than a button, and this is the trade-off that hides from the file listing.
Why Won't Some PDFs Compress at All?
The section above covers files that won't shrink because there's nothing left to squeeze. But there's a second group, and these ones fail differently. The tool doesn't produce a smaller file. It refuses, or throws an error, or just does nothing, and you're left wondering whether the site is broken.
Usually it's the file, and usually it's locked.
Your PDF is encrypted. If the document needs a password to open, a compressor genuinely cannot read the content streams to re-encode them. The qpdf documentation on PDF encryption puts the requirement plainly: you need to know at least one of the user or owner password to retrieve the encryption key. No key, no access to the images inside, no compression. That's the format working exactly as designed, not a failure of the tool.
Your PDF carries permission restrictions. This one is stranger, because the file opens fine and still won't process. Some PDFs have no open password but do carry flags saying modification isn't allowed, and a well-behaved tool respects that and stops.
Worth knowing what those flags actually are, though. The qpdf documentation is direct about it, noting that the security of the restrictions placed on PDF files is solely enforced by the software, and that because both passwords recover the same single encryption key there is fundamentally no way to prevent an application from disregarding the restrictions on a file. So a tool that refuses your file is being polite rather than technically incapable. If the document is yours, the fix is to remove the restriction at the source where you created it, rather than hunting for a tool willing to ignore a flag. If it isn't yours, the flag is telling you something and going around it is a decision, not an accident. For the formal definitions behind any of this, the PDF specification published by Adobe as ISO 32000-1 is the primary source.
And one lever people skip entirely. If your file is a scan of black text on white paper that was captured in color, converting it to greyscale before compressing often does more than any quality slider. A color scan carries three channels of information for a page that only ever needed one, and you're paying for all three on every page.
What Else Do People Ask?
How do you compress a PDF for free?
Open a free browser-based PDF compressor, drop in your file, pick a compression level, and download the smaller version. Good tools do this locally, so your PDF never leaves your device. The whole process takes under a minute, needs no signup, and usually cuts the size noticeably, more on text-heavy files than on already-compressed scans.
Does compressing a PDF reduce quality?
It depends on the method. Lossless compression strips unused data and reorganises the file without touching anything you can see, so quality stays identical. Lossy compression downsamples and re-encodes images to save more space, which can soften photos if you push it hard. For text and line art, you can compress a lot with no visible change.
What is the difference between lossy and lossless PDF compression?
Lossless compression keeps every bit of the original and shrinks the file by removing redundancy, typically saving 20 to 40 percent. Lossy compression permanently removes some image detail to save much more, often 50 to 90 percent, at the cost of some picture quality. Most tools blend both, compressing images lossily while keeping text lossless.
How small can you make a PDF for email?
Aim under about 10MB to be safe everywhere, since limits vary. Gmail and Outlook.com both cap attachments at 25MB, but email encodes files in Base64, which adds roughly 37 percent overhead, so a 20MB file actually travels as around 27MB and bounces off a limit it looks like it should clear. Corporate Exchange servers are often set lower still, commonly 20MB. Compressing to 10MB or less clears almost every provider without bouncing.
Why won't my scanned PDF compress?
Because scanned pages are stored as images that are usually already compressed inside the file, so there's little redundant data left to strip. A browser tool that only optimises structure will show almost no reduction. To shrink a scan you need a tool that re-encodes the images themselves, which means accepting some quality loss or using a server-side compressor.
Sources: PDF Association on the ISO 32000-2 PDF standard and its compression filters; United States Library of Congress digital format description for PDF 2.0; Google Gmail help on the 25MB attachment limit; United States National Archives transfer guidance format tables, updated August 2025, for the accepted compression methods and the 20 to 1 visually lossless ceiling on digitised permanent records, for PDF/A-1 and PDF/A-2 being preferred over plain PDF for born-digital text, and for the requirement that all fonts including the base 14 be embedded; 36 CFR Part 1236 Subpart E, in force since June 2023, for the 300 pixels per inch minimum when digitising permanent federal records; and David Kriesel's 2013 research documenting JBIG2 Pattern Matching and Substitution altering digits on Xerox WorkCentre and ColorQube scanners. On fonts: Prepressure's reference on PDF font embedding and subsetting, for the definition of a subset and its two trade-offs on editing and merging; and Google's Noto project for the figure that Noto Sans CJK carries 65,535 glyphs, the maximum a single OpenType font can hold. All linked above. We give the glyph count rather than a megabyte figure for CJK fonts, because the file sizes quoted around the web vary by build and we could not verify one from the project itself. The resolution guidance here is general advice for everyday files, not a compliance standard for your own records. Note that NARA's guidance does not name JBIG2 anywhere, so we cite Kriesel's own primary research for the substitution behaviour rather than attributing it to a regulator. NARA states the base 14 font requirement in its rules for PDF case file collections rather than as a blanket rule covering every record type, and we've described it that way rather than overstating its scope; we cite it because it puts the archival principle in plain words, and because the PDF/A formats NARA lists as preferred require embedded fonts by design. An earlier version of this article linked NARA's older PDF records page for the lossy compression point. That page now carries a notice saying it has been superseded and is no longer accurate, so this version cites the current transfer guidance tables instead.
Related Articles
Related Tools