Skip to content
FileSlimmer
Tools
Guides

Why did my PDF barely get smaller?

Last reviewed

I ran my PDF through a compressor and it went from 4.1 MB to 4.0 MB. Is the tool broken?

4 min read · reviewed August 17, 2026

A PDF is a filing cabinet

The mental model that causes the disappointment is thinking of a PDF as a picture of a document. It is not. A PDF, as defined by ISO 32000, is a container of numbered objects: page trees, content streams full of drawing operators, embedded font programs, image XObjects, form fields, annotations, outlines, and a cross-reference table that says where each object begins.

Most of those objects are already compressed. Content streams are usually Flate-encoded, which is the same algorithm a ZIP file uses. Embedded photographs are usually stored as DCT streams, which is JPEG data passed through unchanged. So when a general-purpose compressor looks at a PDF, it is looking at a box of things that have each already been compressed once.

Find out where the bytes are before you compress anything

The single most useful diagnostic is knowing what kind of document you have, because it predicts the result almost completely.

What dominates the size of a PDF, by document type
Document typeWhere the bytes areWhat a structural save can do
Text report from a word processorEmbedded font programs and small content streamsVery little; the content is already dense
Scanned pagesOne large image per pageAlmost nothing without touching the images
Presentation exportEmbedded photographs and gradientsSome, if images are duplicated across slides
Engineering drawingLong vector content streamsSome, by re-encoding streams and removing unused objects
Form with attachmentsEmbedded files, JavaScript, annotation dictionariesA lot, if the extras are removable

What a structural save actually removes

A structural save rewrites the document without changing what any page looks like. It can drop objects nothing references any more, which accumulate every time a document is edited and saved incrementally. It can compress the cross-reference table and pack small objects into object streams. It can drop document metadata, thumbnails, and the editing history that some producers leave behind.

On a document that has been edited and re-saved many times, that can be a large win. On a document exported cleanly, once, from a word processor, there is nothing to collect: the file is already close to the smallest form of itself, and a hundred kilobytes off four megabytes is exactly the result to expect.

Why a text document resists

In a text PDF, a large share of the file is usually embedded fonts. A font program is a compact binary containing glyph outlines and hinting instructions, and it is there because the document must render identically on a machine that does not have that typeface installed. You cannot compress it away without either subsetting it further or removing it, and removing it changes how the page looks.

The text itself is tiny. A hundred pages of prose is a few hundred kilobytes of characters before compression. If your text PDF is large, look at the fonts and at any images hiding behind the text — a logo repeated in a header, a background watermark — rather than at the words.

Why a scan collapses

A scanned page is one big photograph of a piece of paper, often captured at 300 dots per inch or more. An A4 page at 300 DPI is roughly 2480 × 3508 pixels — nearly nine million of them, for a page whose information content is a few kilobytes of text. That is why scans respond so dramatically to rasterised compression: reducing the render resolution and re-encoding the page images attacks the part of the file that is actually large.

It is also why that mode is destructive. Once a page is a picture, the searchable text layer, the links, the form fields, the annotations, the tagged accessibility structure and any digital signature are gone. FileSlimmer treats rasterisation as a separate, clearly-labelled mode with that warning attached, rather than as a silent fallback when the structural save disappoints.

Why nobody can promise you a number

A target size for a PDF is a search over render resolution and image quality, and the search has a floor: below some resolution, the text on a scan stops being readable. Whether a particular document reaches a particular budget above that floor depends on the page count, the ink coverage, the amount of photographic content and the scanner's noise. Two documents with the same page count can end up at very different sizes.

The honest behaviour is to stop at the floor and report where the search stopped, which is what FileSlimmer's PDF target presets do. A tool that always hits the number you typed is either lying about the number or destroying the document to reach it.

Before you blame the tool

  • Check whether the PDF is a scan. Try selecting a line of text: if the cursor will not select anything, every page is an image.
  • Check the page count against the size. A 40 MB file with three pages is image-heavy; a 40 MB file with 900 pages may be entirely reasonable.
  • Check whether the document is encrypted. A password-protected PDF cannot be restructured at all until it is unlocked.
  • Check whether it is signed. Rewriting a signed document invalidates the signature, and that is usually a worse outcome than a large file.
  • Check what you actually need. Splitting out the six pages you have to send often beats compressing all ninety.

Tools for this

Sources

All FileSlimmer guides