A lossless PDF compression pass rewrites the file's internal structure rather than re-encoding the images inside it. We measured what that actually saves across eight test PDFs. Three of them came out larger in at least one mode, and the worst regression was 20.8 percent.
The saving depends on how loose the input was. Uncompressed text streams shrank by 84 to 90 percent, while a file whose streams were already Flate-compressed shrank by 31.6 percent in standard mode and grew by 3.0 percent in maximum mode. A second pass changed nothing on any sample, and two files with the same content but different structure produced byte-identical output. None of this means the engine is faulty; it means a size claim needs a measurement behind it. The full method, the table and the limits of the test are below.
Every PDF tool has a Compress button, and almost all of them report success. That word covers at least two very different operations, and one of them can hand you back a bigger file. We wanted numbers instead of a promise, so we built eight PDFs with deliberately different internal structures and ran each one through qpdf 11.7.0 using the same arguments our own tool passes.
Compress PDF
The same qpdf pass described here, compiled to WebAssembly and run inside the tab. Pick the level, then read the before and after sizes yourself.
What a "compress PDF" button actually does
The first thing worth knowing is that this is usually not image re-compression. It is a structural rewrite. A PDF holds its pages as a collection of objects and streams, and the pass unpacks them, removes duplication, compresses the streams that were stored raw, and writes a new file.
qpdf, the engine inside our tool, does that with this call in standard mode:
--object-streams=generate --compress-streams=y --recompress-flate --compression-level=6
and this one in maximum mode, which adds a higher level and linearisation for fast web viewing:
--object-streams=generate --compress-streams=y --recompress-flate --compression-level=9 --linearize
Notice what is not in either call: anything that touches the embedded images. A structural pass does not decode a JPEG that already sits inside a PDF, so it cannot meaningfully shrink a scanned or photo-heavy document. What it can do is remove the inefficiency that accumulated while the file was being built. That is why the saving depends on the input rather than on how hard the tool tries.
How the measurement was run
Eight samples were generated to isolate one structural property each: a minimal one-page file, a three-page file, the same twenty-page document with compressed and uncompressed streams, a hundred-page file with many objects, a one-page file padded with a junk comment, a file with an embedded JPEG, and a file with a small uncompressed XMP packet. Every file was then processed with the arguments above, twice in the case of standard mode, and the byte count and SHA-256 hash were recorded each time.
- Generate the samples by script, not by hand. That makes each one reproducible and keeps the test from depending on whichever documents happened to be on hand.
- Use the tool's real arguments. The engine is called with exactly the flags the browser tool passes, so the numbers describe the product rather than a similar-looking experiment.
- Record bytes and hashes, not just a percentage. A hash answers a second question the size cannot: whether two differently built inputs end up as the same file.
- Run everything a second time. Repeating the operation tests whether it converges, which matters if an interface invites you to press the button again.
The numbers
Standard mode is level 6 without linearisation. Maximum is level 9 with linearisation. Negative percentages are smaller files.
| Sample | Input (B) | Standard (B) | Change | Maximum (B) | Change |
|---|---|---|---|---|---|
| 01 minimal, 1 page | 3,392 | 872 | -74.3% | 1,637 | -51.7% |
| 02 minimal, 3 pages | 9,520 | 1,475 | -84.5% | 2,513 | -73.6% |
| 03 uncompressed text, 20 pages | 62,145 | 6,590 | -89.4% | 9,916 | -84.0% |
| 04 compressed text, 20 pages | 9,629 | 6,590 | -31.6% | 9,916 | +3.0% |
| 05 many objects, 100 pages | 311,281 | 30,964 | -90.1% | 45,110 | -85.5% |
| 06 loose structure, 1 page | 7,489 | 872 | -88.4% | 1,637 | -78.1% |
| 07 embedded JPEG image | 67,974 | 68,124 | +0.2% | 68,863 | +1.3% |
| 08 uncompressed XMP | 4,151 | 4,235 | +2.0% | 5,013 | +20.8% |
Four things the numbers say
The gain comes from the input being loose, not from the tool being clever
Files whose text streams were stored uncompressed shrank by 84 to 90 percent. A file whose streams were already Flate-compressed shrank by 31.6 percent in standard mode and then grew by 3.0 percent in maximum mode, because level 9 plus linearisation added structure the file did not need. More effort is not the same as a smaller file, and the limit is the input's structure, not the effort applied to it.
Three of the eight samples got bigger
The worst case was a 20.8 percent increase on a file carrying a small uncompressed XMP packet, pushed through maximum mode. The sample with an embedded JPEG grew in both modes, because there was nothing structural to win and the rewrite itself costs bytes. This is the part that a success message hides: a compressor that always reports success is not measuring anything, and on these inputs the honest output is a slightly larger file.
The operation is idempotent on its own output
Every sample was run a second time through the standard path, and every second-pass change was 0.0 percent. Once a file has been rewritten, rewriting it again finds nothing. If a tool appears to keep finding fresh savings each time you press the button, that is a reason to look at what it is actually doing rather than a sign of thoroughness.
The rewrite normalises, it does not merely trim
Samples 01 and 06 are the same one-page document, except that 06 carries four kilobytes of junk comment. After the rewrite they produce the same SHA-256 hash. Samples 03 and 04 are the same twenty-page document with compressed and uncompressed streams, and they converge the same way. The engine is rebuilding the file into a canonical form rather than shaving whatever it happens to find, which is also why a second pass has nothing left to do.
How to tell whether a compressor measured anything
Two checks take about a minute each and work with any tool.
- Ask for the number. A message that says Compressed without a byte count before and after has told you nothing. The cases worth watching are small files and already-well-made files, where the honest answer is often that it grew.
- Try a file that is already tight. An export from a modern office suite is usually already compressed. If a tool still claims a large saving on it, it is either re-encoding your images, which is a quality decision rather than a lossless one, or it is guessing.
Both checks point at the same idea. A structural optimiser has no target size to aim for, which is why our own tool has no field for one and why the honest description of what it can do depends on the document in front of it.
Where this test has limits
The numbers above are real and they are narrow. Being clear about the gaps is the point of publishing them.
- The samples are generated, not collected. They isolate structural properties on purpose, so they are not a random sample of the PDFs people actually have.
- Eight is a small set. Read the percentages as evidence that a range exists, not as an average for all PDFs.
- Images are represented by one sample. A single embedded JPEG stands in for a whole class of documents. Deeper nesting of images, fonts and form fields was not tested.
- The engine never re-encodes images here. That is a property of qpdf's structural pass, not something this dataset proves about other compressors.
- Versions matter. These results belong to qpdf 11.7.0. If a later version changes its optimisation, the numbers change with it and the method has to be run again.
Frequently asked questions
If you want the wider picture of what a browser tool can and cannot do with a document, the guide on compressing a PDF without uploading it covers the engine and the file path, and the question of target sizes explains why a lossless optimiser has no field for one.