PolyPress

Handrail · a Mac app that writes R for you to run

It writes R into
your script.

Handrail is a Mac app that writes R into your RStudio script. Then it gets out of the way.

It is for people who know what they want to do with their data but not how to say it in R. You pick the step from a menu; it writes the real code, with your plain-English sentence above it as a comment. It is not a no-code tool and not a replacement for RStudio — it is a scaffold you are meant to stop needing.

You point it at an RStudio project, pick a script and a data file, build a step from menus, and press Add to script. Real dplyr goes into that script. RStudio notices the file changed on disk and reloads it, so the line appears in front of you — and you run it there, with Cmd-Enter, in RStudio, like anybody else.

Handrail does not run R. It does not display your data as the source of truth, draw charts, hold results, manage packages or open a console. Its entire job is the gap between “I know what I want to do” and “I know how to write it in R”. Nothing else.

01

What it actually writes

Choosing “Keep only the rows where age is more than 65” appends:

# Keep only the rows where age is more than 65
data <- data |>
  filter(age > 65)

Above every generated line sits the plain-English sentence you chose, as a comment. That comment is the only part of the app that survives the app being closed — which is the whole point. Next time you open the script, the English and the R are side by side.

The first step also writes the preamble, once, and only if the script doesn't already read the data:

# Built with Handrail. This is ordinary R — it runs with or without the app.
library(dplyr)

data <- read.csv("data/people.csv", stringsAsFactors = FALSE)

The path is relative to the project, so the script still runs on another machine. If the data file lives outside the project, a comment says so and the app warns you in the form.

An answer is written as two statements — the table gets a name, then the name sits on its own line so running it shows you something:

# Count how many rows have each value of region, commonest first, with percentages
region_counts <- people |>
  count(region, name = "n") |>
  mutate(percent = round(100 * n / sum(n), 1)) |>
  arrange(desc(n))

region_counts
02

The mistake it exists to correct

Handrail began on 2026-08-20, after two earlier attempts were scrapped. The first was an R package of plain-language verbs. The second was a Shiny GUI that had, without anyone deciding to, grown into a whole analysis environment.

Both failed the same way, and it is the thing the project now guards against.

Every hour spent building a table view, a chart picker or a console is an hour spent building a worse RStudio.

The test for any proposed feature: does RStudio already do this? If yes, the answer is no. Viewing data, running code, plotting, debugging, installing packages, managing the project — all RStudio's.

03

The four rules

Enforced in code and pinned by tests. These are the design, not guidelines.

  1. Append, never rewrite

    Whatever is above the line we add belongs to the user — including lines we wrote and they then changed. ScriptWriter only ever appends, and “take that back out” only works if the file still ends with exactly the text we added. If the user has typed since, it refuses and points them at Undo in RStudio. An app that reformats someone's file is one they stop trusting, and this one is trusted with the only copy.

  2. One action, one idea — and every function is one you can look up

    Never a helper only this app knows. The mechanical test: no generated line may call a function that isn't documented in R or in a named package. ggplot() + geom_bar() + labs() passes; a handrail_describe() helper fails.

  3. Nothing is hidden

    na.rm = TRUE is written out rather than relied on, because the next thing you read about mean() will say the default is FALSE. fixed = TRUE is written out on grepl because without it "1.5" also matches "125". row.names = FALSE is spelled out on write.csv because R's default of TRUE is the classic reason a CSV arrives looking wrong.

  4. Half an action never reaches the script

    Code generation returns nil until there is enough to make legal R, and the Add button stays disabled until then.

04

The thirty-three actions

Grouped by intent, not by dplyr verb — grouping by verb makes a glossary, and a glossary is what this app exists to avoid. Each is searchable by the word a beginner would actually type: “missing” finds Drop rows with gaps in them; “xlsx” and “spreadsheet” both find the Excel one.

Fewer rows

  • Keep only some rows
  • Drop rows with gaps in them
  • Keep the top few
  • Drop repeated rows

Boil it down

  • Count, average or total, for each…

Order and columns

  • Sort the rows
  • Keep only some columns
  • Add a column worked out from others
  • Rename a column
  • Round numbers off
  • Sort numbers into bands

Fix a column that came in wrong

  • Fill in the gaps
  • Treat a value as missing
  • Treat a column as numbers
  • Treat a column as dates
  • Treat a column as text
  • Pull the year or month out of a date

Get an answer

  • Count how many of each
  • Describe a number
  • See how much is missing
  • Find rows that repeat

Save it to a file

  • Save this as a CSV
  • Save this to open again in R
  • Save this as an Excel file
  • Save with a different separator
  • Save for Stata, SPSS or SAS
  • Save a summary table, not the rows

Look at it

  • Show the first few rows
  • List every column and what is in it
  • Show the range and average of each column
  • Cross one column against another
  • Open it in RStudio's viewer
  • Say how many rows are left
05

Speed, measured not asserted

The app sits beside RStudio all day. It has to cost nothing. There is a benchmark executable that prints size, seconds, bytes actually read, and peak memory.

A real 278 MB, 209-column survey file.
BeforeAfter
Time to load> 120 s0.04 s
Bytes read278 MB1.74 MB
Process memory> 1.1 GB~103 MB
The two operations that deliberately read more than the head, same file. The twentieth page costs what the first one does — pages resume from a byte offset, so paging doesn't degrade as you go down a file.
OperationCostPeak memory
Count every row of one column (96,539 rows)1.09 s21 MB
One viewer page, 500 rows × 12 of 209 columns6 ms20 MB
The twentieth viewer pagealso 6 ms20 MB

The five mistakes that caused the “before” column

All of them the standard way to get this wrong in Swift, and all pinned by tests now.

  1. Reading the whole file

    Data(contentsOf:) plus String(data:encoding:) holds a 278 MB file twice over, to keep 400 rows. The reader now streams 256 KB at a time until it has enough line breaks, then trims back to the last newline — always a safe cut in UTF-8, since 0x0A never appears inside a multi-byte sequence.

  2. Character

    Swift's Character is a grapheme cluster, so iterating a String runs Unicode segmentation per step. Splitting on , and \n needs none of it. The CSV parser works on UTF-8 bytes and decodes only finished fields.

  3. Work on the render path

    A SwiftUI body runs on every keystroke, so anything in one runs hundreds of times. Three were: a 209-entry type dictionary rebuilt per read, a column-list filter run per picker per render, and a directory scan called from inside a view body.

  4. FileHandle.read(upToCount:) returns autoreleased Data

    Reading a file in a loop without an autoreleasepool inside the loop keeps every chunk alive until the function returns: counting one column of the 278 MB file peaked at 304 MB, and 21 MB with the pool. Same time either way — which is why nothing but a measurement would have found it.

  5. The view tree

    A plain VStack in a ScrollView is eager, and every column row carried its own Menu: 209 columns produced thousands of views at once and SwiftUI's AttributeGraph aborted the process with SIGABRT. Now LazyVStack, one popover built only for the open column, and a filter box past a dozen columns.

06

The most valuable test runs R

225 assertions. The best of them writes a real CSV and a real script, runs it with Rscript --vanilla, and compares the answer. Every action the app offers is checked this way. Text-comparison tests can't catch a wrong quote or a column name R can't see; this does.

Quoting is the whole ballgame. filter(age == "65") is legal R that silently keeps nothing. The literal generator decides by the column's sniffed type, and deliberately leaves a non-numeric value quoted even in a number column, so R errors rather than lying.

Two tests exist specifically to stop the answer actions going wrong: every .result action must leave nrow(data) unchanged, and every action in the “Get an answer” group must return a non-nil caution — without that test a new answer action ships silent, and these are exactly the ones whose wrong answers look right.

07

Status, and one unresolved problem

Under git since 2026-08-21. Working today: project, script and data selection, thirty-three actions, the column menu, the searchable palette, the code preview, the cautions, appending, taking the last one back, opening the script in RStudio, the menu-bar panel, and the data viewer window. About 6,200 lines of Swift. Requires macOS 14 or newer. No R packages, no installer, no account.

The evidence that “getting an answer out” was the gap: across 81 real R scripts on the author's machine, every project ends in a Table 1 with p-values, a gt or flextable table, and a figure. The first twenty-nine actions all prepared data; none produced a result.

The unresolved problem, and it is a real one

A naive t-test or chi-squared on survey-weighted data is wrong. A dataset like NEDS needs the design — strata, clusters and weights — which means survey::svydesign with svyttest / svychisq, not base R. So when a weight column is set, Handrail must either refuse to write p-values and say why, or generate real survey code.

Refusing with an explanation is the honest floor; generating survey code is the useful answer, and needs the strata and cluster columns named too. This has to be decided before the p-value generator is written, because writing naive weighted tests would be the worst bug this app could ship — the numbers would look fine.

Table 1 is specced and compiles, but nothing is wired to it. Three decisions are already taken: plain dplyr and base R rather than gtsummary or tableone, because both choose a test per variable by their own rules and that is the one decision that isn't theirs to make; tests are suggested and confirmed, never applied, with the reason shown beside the suggestion and the test's name going into the table beside the p-value; and weighting is an action, not a table option, so every summarising step after it generates weighted arithmetic explicitly.