What Lossless PDF Compression Actually Does to File Size

A lossless PDF compression pass rewrites the file's internal structure rather than re-encoding the images inside it. We measured what that actually saves across eight test PDFs. Three of them came out larger in at least one mode, and the worst regression was 20.8 percent.

The saving depends on how loose the input was. Uncompressed text streams shrank by 84 to 90 percent, while a file whose streams were already Flate-compressed shrank by 31.6 percent in standard mode and grew by 3.0 percent in maximum mode. A second pass changed nothing on any sample, and two files with the same content but different structure produced byte-identical output. None of this means the engine is faulty; it means a size claim needs a measurement behind it. The full method, the table and the limits of the test are below.

Every PDF tool has a Compress button, and almost all of them report success. That word covers at least two very different operations, and one of them can hand you back a bigger file. We wanted numbers instead of a promise, so we built eight PDFs with deliberately different internal structures and ran each one through qpdf 11.7.0 using the same arguments our own tool passes.

Compress PDF

The same qpdf pass described here, compiled to WebAssembly and run inside the tab. Pick the level, then read the before and after sizes yourself.

Open Compress PDF

What a "compress PDF" button actually does

The first thing worth knowing is that this is usually not image re-compression. It is a structural rewrite. A PDF holds its pages as a collection of objects and streams, and the pass unpacks them, removes duplication, compresses the streams that were stored raw, and writes a new file.

qpdf, the engine inside our tool, does that with this call in standard mode:

--object-streams=generate --compress-streams=y --recompress-flate --compression-level=6

and this one in maximum mode, which adds a higher level and linearisation for fast web viewing:

--object-streams=generate --compress-streams=y --recompress-flate --compression-level=9 --linearize

Notice what is not in either call: anything that touches the embedded images. A structural pass does not decode a JPEG that already sits inside a PDF, so it cannot meaningfully shrink a scanned or photo-heavy document. What it can do is remove the inefficiency that accumulated while the file was being built. That is why the saving depends on the input rather than on how hard the tool tries.

How the measurement was run

Eight samples were generated to isolate one structural property each: a minimal one-page file, a three-page file, the same twenty-page document with compressed and uncompressed streams, a hundred-page file with many objects, a one-page file padded with a junk comment, a file with an embedded JPEG, and a file with a small uncompressed XMP packet. Every file was then processed with the arguments above, twice in the case of standard mode, and the byte count and SHA-256 hash were recorded each time.

  1. Generate the samples by script, not by hand. That makes each one reproducible and keeps the test from depending on whichever documents happened to be on hand.
  2. Use the tool's real arguments. The engine is called with exactly the flags the browser tool passes, so the numbers describe the product rather than a similar-looking experiment.
  3. Record bytes and hashes, not just a percentage. A hash answers a second question the size cannot: whether two differently built inputs end up as the same file.
  4. Run everything a second time. Repeating the operation tests whether it converges, which matters if an interface invites you to press the button again.

The numbers

Standard mode is level 6 without linearisation. Maximum is level 9 with linearisation. Negative percentages are smaller files.

SampleInput (B)Standard (B)ChangeMaximum (B)Change
01 minimal, 1 page3,392872-74.3%1,637-51.7%
02 minimal, 3 pages9,5201,475-84.5%2,513-73.6%
03 uncompressed text, 20 pages62,1456,590-89.4%9,916-84.0%
04 compressed text, 20 pages9,6296,590-31.6%9,916+3.0%
05 many objects, 100 pages311,28130,964-90.1%45,110-85.5%
06 loose structure, 1 page7,489872-88.4%1,637-78.1%
07 embedded JPEG image67,97468,124+0.2%68,863+1.3%
08 uncompressed XMP4,1514,235+2.0%5,013+20.8%

Four things the numbers say

The gain comes from the input being loose, not from the tool being clever

Files whose text streams were stored uncompressed shrank by 84 to 90 percent. A file whose streams were already Flate-compressed shrank by 31.6 percent in standard mode and then grew by 3.0 percent in maximum mode, because level 9 plus linearisation added structure the file did not need. More effort is not the same as a smaller file, and the limit is the input's structure, not the effort applied to it.

Three of the eight samples got bigger

The worst case was a 20.8 percent increase on a file carrying a small uncompressed XMP packet, pushed through maximum mode. The sample with an embedded JPEG grew in both modes, because there was nothing structural to win and the rewrite itself costs bytes. This is the part that a success message hides: a compressor that always reports success is not measuring anything, and on these inputs the honest output is a slightly larger file.

The operation is idempotent on its own output

Every sample was run a second time through the standard path, and every second-pass change was 0.0 percent. Once a file has been rewritten, rewriting it again finds nothing. If a tool appears to keep finding fresh savings each time you press the button, that is a reason to look at what it is actually doing rather than a sign of thoroughness.

The rewrite normalises, it does not merely trim

Samples 01 and 06 are the same one-page document, except that 06 carries four kilobytes of junk comment. After the rewrite they produce the same SHA-256 hash. Samples 03 and 04 are the same twenty-page document with compressed and uncompressed streams, and they converge the same way. The engine is rebuilding the file into a canonical form rather than shaving whatever it happens to find, which is also why a second pass has nothing left to do.

How to tell whether a compressor measured anything

Two checks take about a minute each and work with any tool.

  1. Ask for the number. A message that says Compressed without a byte count before and after has told you nothing. The cases worth watching are small files and already-well-made files, where the honest answer is often that it grew.
  2. Try a file that is already tight. An export from a modern office suite is usually already compressed. If a tool still claims a large saving on it, it is either re-encoding your images, which is a quality decision rather than a lossless one, or it is guessing.

Both checks point at the same idea. A structural optimiser has no target size to aim for, which is why our own tool has no field for one and why the honest description of what it can do depends on the document in front of it.

Where this test has limits

The numbers above are real and they are narrow. Being clear about the gaps is the point of publishing them.

  • The samples are generated, not collected. They isolate structural properties on purpose, so they are not a random sample of the PDFs people actually have.
  • Eight is a small set. Read the percentages as evidence that a range exists, not as an average for all PDFs.
  • Images are represented by one sample. A single embedded JPEG stands in for a whole class of documents. Deeper nesting of images, fonts and form fields was not tested.
  • The engine never re-encodes images here. That is a property of qpdf's structural pass, not something this dataset proves about other compressors.
  • Versions matter. These results belong to qpdf 11.7.0. If a later version changes its optimisation, the numbers change with it and the method has to be run again.

Frequently asked questions

It rewrites the file's internal structure rather than re-encoding the pictures inside it. A PDF stores its pages as a set of objects and streams, and a compression pass unpacks those objects, deduplicates what repeats, compresses the streams that were stored uncompressed, and writes a new file. The engine we measured, qpdf, does exactly that when called with object streams enabled, stream compression on and Flate recompression at level 6. Nothing in that call touches the embedded images, so a JPEG that is already inside the document is copied across unchanged. That is what makes the operation lossless: the visible content is the same after the rewrite, and the only thing that changed is how the file is organised on the inside. It is also why the saving is so uneven, and why a photo-heavy document can barely move at all.
Yes, and it happened to three of our eight samples. The largest increase was 20.8 percent, on a file carrying a small uncompressed XMP metadata packet pushed through the maximum mode, which adds level 9 compression and linearisation. A second sample with an embedded JPEG grew in both modes, by 0.2 percent in standard mode and 1.3 percent in maximum mode, because there was almost nothing structural to gain and the rewrite still costs bytes. The pattern behind all three cases is the same: when the input is already tight, the overhead of rebuilding it is larger than the small saving the rebuild finds. A tool that reports success on every file without saying how much it saved cannot be checked, and it will happily return a bigger file.
Not in the mode we measured, because the images are never touched. The pass rewrites object structure and recompresses the streams that hold page content, fonts and metadata. An image that is already encoded as JPEG inside the PDF is stored as a stream of compressed bytes, and those bytes are copied, not decoded and re-encoded. This is a property of the engine rather than a promise about every tool. Some compressors offer a second, lossy mode that downsamples embedded images to a chosen DPI and re-encodes them as JPEG, which does reduce quality and does shrink photo-heavy files dramatically. Those two operations are often presented under one button. The way to tell them apart is the output size on an image-heavy document: a lossless pass barely moves it, while a lossy pass can cut it by half.
Because the result tracks the arguments, the engine version and the input, and those differ between tools. We measured one engine at one version with two argument sets and got a spread from 90 percent smaller to 20.8 percent larger on different inputs. Change the compression level, add linearisation, or enable image downsampling and the numbers move again. The input matters just as much: in our set, files whose text streams were stored uncompressed shrank by 84 to 90 percent, while a file whose streams were already Flate-compressed shrank by 31.6 percent in standard mode and grew in maximum mode. Two tools can both be honest and still report different numbers for the same document, which is why a result quoted without the arguments and the input it came from is not really a measurement.
Compare the byte count before and after, and treat a missing number as the answer. A tool that prints a success message without showing both sizes has told you nothing, and the interesting cases are exactly the ones where it would have to admit a larger file. On a computer, the file's properties or a directory listing gives you the exact size in bytes. Then read the direction, not just the percentage: a two percent change either way is within the range where a rewrite can go against you. If the tool offers a quality or level setting, run the same file twice at different settings so you can see whether the numbers move at all. Our own tests are in the same spirit, and the exact qpdf arguments are printed in this article so the same run can be repeated with the same engine.

If you want the wider picture of what a browser tool can and cannot do with a document, the guide on compressing a PDF without uploading it covers the engine and the file path, and the question of target sizes explains why a lossless optimiser has no field for one.

Check the numbers on your own file

Free, private and unlimited. No account needed.

Open Compress PDF