Skip to content

fix: audit 2026-10-09 - rclone rc on a private socket, fail-closed remote check, lost process output - #21

Merged
mbeczynski merged 3 commits into
mainfrom
fix/audyt-2026-10-09
Oct 9, 2026
Merged

mbeczynski merged 3 commits into
mainfrom
fix/audyt-2026-10-09

Conversation

@mbeczynski

Copy link
Copy Markdown
Contributor

Fixes the findings of the 9 Oct 2026 audit: K1, W1-W3, S1-S5 and the cheap LOW items. Each fix comes with a regression test. Every test was checked against the old behaviour: I reverted each fix in turn and confirmed its tests fail (results below).

K1 (critical): rclone remote control on a private socket

What was wrong. rclone mount listened on 127.0.0.1:5572 with --rc-no-auth, and the mount runs with --drive-use-trash=false. Loopback does not keep web pages out:

  • A page can send a form POST to it without a CORS preflight.
  • rclone does not check Origin.
  • So any page open in the browser could call operations/purge on gdrive:CloudMachine/<folder>, and the delete would be permanent.

The audit confirmed this against the running mount: config/listremotes answered a form POST sent with Origin: https://evil.example.

Fix

  • The mount now listens on --rc-addr unix://~/.cloudmachine/run/rc.sock, in a directory set to 0700 (re-applied on every start).
    • A browser cannot reach a unix socket at all.
    • Other users cannot enter the directory.
    • Checked by hand on rclone 1.75.1 with a throwaway rclone rcd in /tmp: the socket works, and the client reaches it with rclone rc --unix-socket.
  • One way in. Every caller goes through DriveBufferService.rc(...). That covers vfs/stats, vfs/queue and vfs/queue-set-expiry, used by the GUI, drive-status, buffer-guard, backup-health and detach. A test fails if any source file builds its own ["rc", ...] call. There are no rc calls in scripts or docs.
  • Stale socket. rclone does not remove its socket after a crash or kill -9, and the next start then fails with bind: address already in use (checked on 1.75.1). Under KeepAlive the mount would never come back. So prepare() now:
    • removes a socket that nobody answers on;
    • leaves alone a socket that answers, because it belongs to a running rclone. The new start then fails the same way a second mount on the same TCP port used to.
  • Path limit. If the socket path is longer than macOS allows (103 bytes), prepare() stops the start with a clear message. Without that check, rclone fails with invalid argument.
  • --rc-no-auth stays, on purpose. On a socket nothing gains from a password:
    • vfs/queue-set-expiry, which every detach needs, requires either a password or --rc-no-auth.
    • The only callers that can still reach the socket are processes of the same user. Those can read rclone.conf, token included, directly.
    • A per-start password would be one more way for clients to lose the queue after a restart.
    • The test checks the address instead: the only --rc-addr is the socket, and there is no 127.0.0.1, localhost or :5572 anywhere in the arguments.
  • Requirement. The client flag --unix-socket needs rclone ≥ 1.68 (Sept 2024). Every CloudMachine install downloaded a newer one: install-rclone takes the latest, and the project dates from 2026.

Drive trash (--drive-use-trash=true) as a second layer: not enabled.

  • Bands are rewritten in place, so uploads would not fill the trash.
  • Deletions would. The big deletions are deliberate ones (re-creating the image, clearing an old folder). Each would keep hundreds of GB counted against the account for 30 days. On this account, running out of space stops the mount (storageQuotaExceeded).
  • The free-space alarm would also report space that no folder shows, because operations/about counts the trash.
  • With the socket in place, the trash would only protect against a process of the same user deleting through the mount. Such a process can just as well run its own rclone purge --drive-use-trash=false.

The reasoning is written up in docs/design.md and in a comment next to the flag.

W1: isConfigured fails closed

  • RemoteConfigurer.isConfigured returned false on a timeout, a failed start or exit code 1. An unreadable or broken rclone.conf gives exit 1 (checked).
  • It now returns Bool?, matches whole lines (mygdrive: is not gdrive:), and requires succeeded.
  • connect stops on nil with a message, even with --replace-existing.
  • DriveFolder.decide(legacyEvidence: nil) refuses whenever the answer would decide the name.
  • The window keeps its last known answer, instead of offering "Connect Google Drive" on a working installation.

Not done: the audit's extra local evidence for a legacy installation. There is no reliable local trace:

  • the cache directory is named gdrive{hash};
  • agents-version is written at the first launch of any install.

W3: ProcessRunner lost output

Cause. terminationHandler removed the readabilityHandlers before the last data in the pipe was read. Reads also ran on a separate Foundation queue, so a chunk already read could reach the serial queue after the result had been built.

Fix

  • Both pipes are read by DispatchSourceRead on the same serial queue that builds the result.
  • After the process exits, a final non-blocking read collects what is left. It stops at EAGAIN, not at EOF, so a --daemon child that keeps the pipe open cannot hold the result.
  • The reader keeps draining, and discarding, until EOF. A child that inherited the pipe therefore never blocks on a full pipe.
  • The timeout path keeps the same "keep draining" behaviour.

Test. 8 tasks × 500 runs of /bin/cat on a 4.4 kB file must return exactly the file every time. It also checks stderr, and that an inherited pipe (sleep 20 &) does not delay the result.

  • On origin/main: 970 of 4000 calls lost their output.
  • On this branch: 0, in 5 runs.

TimeMachineStatus.output now returns nil (no answer) when the exit code is non-zero or the output is empty. The exception is "No destinations configured", which still counts as an answer, whatever exit code comes with it. Checked here: tmutil status and tmutil destinationinfo both exit 0 on a working system.

W2: false "no Full Disk Access" alarm

  • Watchdog: reads the Time Machine preferences up to 3 times, 5 s apart (BackupHealth.readPreferences).
  • Window: shows "missing" only after 2 failed reads in a row (BackupHealth.ReadConfirmation).
    • The first good read restores it at once.
    • At launch the state starts as "missing", so a real lack of permission still shows straight away.
  • Missing access makes every read fail, so the real alarm is only delayed by about 10 s.

S1-S5 and LOW

  • S1: MachineBudget.limitGB looks the limit up under the folder first, then under the stored machine key (MachineIdentity.storedKey). set-limit moves an old entry instead of duplicating it.
    • On this Mac, after the upgrade, the 3500 GB limit takes effect: the usage measurement (rclone size every 6 h) and the 90%/100% alarms start working.
    • Last measurement: 589 GiB, so no alarm.
  • S2: "Back up now" uses canStartBackup: dependencies, remote, mount, attached image, registered TM, !isBusy. An old backup, an unread queue or the daily limit no longer grey it out. This also fixes N5: the menu now respects isBusy.
  • S3: MachineBudget.standing: no measurement, or one older than 12 h, shows amber; no limit shows neutral. This also fixes N2, the inconsistent tone between the row and the sidebar.
  • S4: an empty queue with erroredFiles > 0 shows "N fragments abandoned - only on this Mac". The queue row takes its tone from uploadTone(uploadState) (N3), and the free-space row uses freeSpaceTone (N2).
  • S5: docs/design.md now gives 327-793 GB per day, with the audit's measurement and its caveat about bands smaller than 32 MiB.
  • N9: actions/checkout, upload-artifact and action-gh-release are pinned to SHAs, with the version in a comment.

Skipped

  • N1 (production on 1.3.4): a deployment state, not code.
  • N4 (remembered section, setup steps only on Overview): a UX decision for the owner.
  • N6 (contrast of the gradient, Reduce Transparency on the sidebar): a visual change to the chosen mockup; it needs the owner's eye.
  • N7 (decimal separator, Date.formatted ignoring CM_LANGUAGE): needs locale plumbing from L10n into every formatter. Not cheap, and not risk-free.
  • N8 (rotating rclone.log while rclone runs): logShowsUploadStalled reads the last 30 minutes of that log. A copy-truncate on a running mount could cut exactly those lines and silence the daily-limit detection. It needs its own design.
  • N10 (test canary in the production log): a production file, outside the scope of code changes.

What happens on brew upgrade from 1.3.4

  1. Homebrew quits the app (uninstall quit:), replaces the bundle and reopens it.
  2. AgentRepair.afterLaunch reloads backup-health, gdrive-attach and buffer-guard, so they run the new code.
  3. gdrive-buffer is not restarted. This is deliberate and unchanged: restarting it drops the mount under Time Machine. The running rclone (on production, pid 1758) keeps listening on 127.0.0.1:5572 with no password.
  4. The new clients find no socket and use the old address (RCTransport.legacyTCP). The queue, the watchdog and detach keep working, exactly as before the upgrade.
  5. drive-status shows: Remote control: OPEN on 127.0.0.1:5572 - the mount was started by an older version ... run prepare-shutdown, then restart the Mac.
  6. K1 is closed at the mount's next start: a restart of the Mac, or a crash of rclone.
    • launchd starts the same cloudmachine-agent mount-drive from the new bundle. The plist does not change.
    • prepare() creates ~/.cloudmachine/run (0700), and rclone starts on the socket.
    • From then on nothing listens on 5572, and the clients use the socket by themselves.
    • No manual configuration step is needed.
    • To close it straight away: cloudmachine-agent prepare-shutdown, then a restart.

The other fixes take effect as soon as the app and the agents run the new code, after step 2.

Tests

  • swift build: OK.
  • swift test --skip BackupImageServiceTests: 359 tests, 0 failures, locally. I skipped BackupImageServiceTests because it takes the production image lock on this Mac. CI runs everything.
  • swift format lint --strict --recursive Sources Tests: clean with the toolchain's swift-format. CI uses the Homebrew one.
  • Each fix reverted in turn, with its tests run (all caught):
Fix reverted Tests
W1 remoteListed back to false on failure 2 failures
W1 decide with nil treated as false 1
W3 answer(from:) ignoring the exit code 3
W2 a single read attempt 4
W2 panel flipping on one failure 2
S1 folder-only lookup 1
S2 button tied to healthy 2
S3 no measurement = ok 2
S4 "Everything uploaded" with abandoned bands 2
K1 TCP address in the mount arguments 2
K1 stale socket not removed 1
W3 ProcessRunner on origin/main 970/4000 lost outputs

Not run: the app itself, launchd, hdiutil, a real rclone mount on the socket. This Mac runs the production backup, so I did not touch any of them. The socket mechanism was checked on a separate rclone rcd.

🤖 Generated with Claude Code

mbeczynski and others added 3 commits October 9, 2026 18:23
…er is no answer

`ProcessRunner` sometimes returned an empty stdout with exit code 0.
`terminationHandler` removed the readability handlers before the last data
in the pipe was read, and a read already done could reach the serial queue
after the result was built. Measured with `/bin/cat` of a 4 kB file: 2 in
5000 calls at rest, 970 in 4000 with eight running at a time. Callers took
"" for an answer - `destinationinfo` for "destination not registered",
`listremotes` for "no remote".

The pipes are now read by dispatch sources on the same serial queue that
builds the result, and the final read after exit is non-blocking, so a
`--daemon` child that keeps the pipe open still cannot hold the result.
The reader keeps draining until EOF, so such a child is never blocked on a
full pipe.

`TimeMachineStatus.output` now looks at the exit code and treats an empty
output as no answer; "No destinations configured" stays an answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…losed when rclone is silent

The mount listened for remote control on 127.0.0.1:5572 with
`--rc-no-auth`. Loopback is not private: a web page can POST a form there
without a CORS preflight and rclone does not check Origin, so any page could
run `operations/purge` on the backup folder - with the Drive trash off.
Checked on the running mount: `config/listremotes` answered such a POST.

It now listens only on `~/.cloudmachine/run/rc.sock`, in a 0700 directory.
A browser cannot reach a unix socket, other users cannot enter the
directory. Every client goes through one function (`DriveBufferService.rc`),
which uses the socket when it exists and the old address only while a mount
from an older version is still running - an upgrade does not restart the
mount. `drive-status` names that state on a new "Remote control" line. A
stale socket left by a crashed rclone is removed before the start (rclone
would fail with "address already in use" forever under KeepAlive); a live
one is left alone.

`RemoteConfigurer.isConfigured` returned false when rclone timed out or
failed (an unreadable rclone.conf exits 1), and the two guards built on it
opened: `configure-remote` could overwrite the working token, and
`DriveFolder` could give a legacy installation a new, empty folder. It is
now `Bool?`; no answer stops both, and the window keeps its last answer
instead of offering "Connect Google Drive".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…imit, misleading rows

- One failed read of the Time Machine preferences during a backup raised
  "most often Full Disk Access is missing" (08.10 19:52). The watchdog now
  reads up to three times, 5 s apart; the window needs two failed reads in
  a row. Missing access fails every read, so the real alarm still comes.
- The 3500 GB limit in machines.json sat under the machine key from before
  the folder became the key, and was looked up only under the folder - no
  measurement, no alarm. Both keys are read now; `set-limit` moves the old
  entry.
- "Back up now" followed `healthy`, so an old backup greyed out the button
  that fixes it. It now needs only somewhere to back up to.
- "Used on Google Drive" was green with no measurement, and a stale one
  passed for current. Amber now; no limit is neutral.
- An empty queue with abandoned bands read "Everything uploaded"; the row
  now counts them. The queue row takes the upload verdict's tone, the free
  space row the card's.
- docs/design.md: the measured upload volume (up to 793 GB in a rolling 24 h
  window), the socket, and why the Drive trash stays off.
- Workflow actions pinned to commit SHAs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mbeczynski
mbeczynski merged commit 4d5c5b0 into main Oct 9, 2026
4 checks passed
@mbeczynski
mbeczynski deleted the fix/audyt-2026-10-09 branch October 9, 2026 17:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant