PolyPress

Two tools for people who work with data

Back to
critical thinking.

Two free, open-source tools for researchers. One makes your data files far smaller without changing a single value. The other writes real R code into your RStudio script, so you can read it, run it, and learn from it.

Both are built on the same idea: you should be able to check what the tool did. Not by being an expert, and not by trusting a black box — but by being able to see the working.

01

The argument

A researcher with a real dataset and a real question now has an enormous amount of help available, and almost none of it is checkable. A model writes the analysis. A library picks the statistical test. A file format decides what your data meant. Each of those is a judgment call, and each is made somewhere the researcher cannot see.

The failure this creates is not that the answer is wrong. It is that the answer looks fine. A cross-tabulation that silently drops the missing rows returns a clean table. A significance test run on survey-weighted data without the design returns a plausible p-value. A column of "1.50" read back as 1.5 round-trips through your pipeline without complaint. Nothing errors. Nobody is alerted.

Both tools on this site are attempts at the same correction, from opposite ends of the workflow: put the judgment somewhere the researcher can see it, and prove the claim rather than asserting it.

Not “trust me.” Not “no code needed.”
Something you can read, check, and eventually stop needing.
02

The two tools

Tool 01 Lossless compression
Python · C · macOS app

Polypress

A lossless compressor for data tables — CSV, Parquet, anything shaped like rows and columns. General-purpose compressors see a table as a stream of bytes; columnar formats compress each column on its own. Polypress is built on the observation that the columns of a real table are not independent of each other, and exploits those relationships explicitly.

It was tested against 500 datasets nobody chose — the most-viewed tables in a public government-data catalog, taken in rank order, with every rejection logged. That design is the point: it is the benchmark you build when you expect to be read by someone hostile.

The result, and the 22 losses Source

500/500 round-trip exactly
478/500 smaller than the best of 17 competitors
1.25× median margin over the next-best tool
Tool 02 Code generation
Swift · macOS 14+

Handrail

A native macOS app that writes R into your RStudio script. You pick a script and a data file, build a step from menus, and press Add to script. Real dplyr goes into that file, with the plain-English sentence you chose sitting above it as a comment. RStudio reloads, and you run it — with Cmd-Enter, in RStudio, like anybody else.

It is not “R without code” and not a no-code tool. Every function it writes is one you can look up in the documentation; nothing is a helper only this app knows. It is a scaffold you are meant to eventually stop needing.

How it writes, and what it refuses

33 actions, grouped by intent
0.04 s to open a 278 MB, 209-column file
225 assertions — the best of them run real R
03

What was measured

Neither project asks to be believed. Both were built under the same rule — a component probe may nominate a choice; only a whole-file measurement decides it — and both keep a written record of what failed.

On Polypress, four separate transforms measured +12.7%, −24.9%, 0.647× and −1.02% in isolation, and three of those four made the whole codec worse. On Handrail, five standard Swift mistakes turned a 120-second file open into 0.04 seconds; one of them cost 304 MB of memory and no amount of reasoning would have found it, because it took the same time either way.

Those negative results are treated as the most valuable part of the documentation, and they are on these pages for the same reason.

500unselected datasets, 3.57 GB of CSV
17competing codecs measured against
22losses, published in full
81real R scripts read to find the gap