Repository navigation
fix: audit 2026-10-09 - rclone rc on a private socket, fail-closed remote check, lost process output - #21
Merged
Conversation
…er is no answer `ProcessRunner` sometimes returned an empty stdout with exit code 0. `terminationHandler` removed the readability handlers before the last data in the pipe was read, and a read already done could reach the serial queue after the result was built. Measured with `/bin/cat` of a 4 kB file: 2 in 5000 calls at rest, 970 in 4000 with eight running at a time. Callers took "" for an answer - `destinationinfo` for "destination not registered", `listremotes` for "no remote". The pipes are now read by dispatch sources on the same serial queue that builds the result, and the final read after exit is non-blocking, so a `--daemon` child that keeps the pipe open still cannot hold the result. The reader keeps draining until EOF, so such a child is never blocked on a full pipe. `TimeMachineStatus.output` now looks at the exit code and treats an empty output as no answer; "No destinations configured" stays an answer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…losed when rclone is silent The mount listened for remote control on 127.0.0.1:5572 with `--rc-no-auth`. Loopback is not private: a web page can POST a form there without a CORS preflight and rclone does not check Origin, so any page could run `operations/purge` on the backup folder - with the Drive trash off. Checked on the running mount: `config/listremotes` answered such a POST. It now listens only on `~/.cloudmachine/run/rc.sock`, in a 0700 directory. A browser cannot reach a unix socket, other users cannot enter the directory. Every client goes through one function (`DriveBufferService.rc`), which uses the socket when it exists and the old address only while a mount from an older version is still running - an upgrade does not restart the mount. `drive-status` names that state on a new "Remote control" line. A stale socket left by a crashed rclone is removed before the start (rclone would fail with "address already in use" forever under KeepAlive); a live one is left alone. `RemoteConfigurer.isConfigured` returned false when rclone timed out or failed (an unreadable rclone.conf exits 1), and the two guards built on it opened: `configure-remote` could overwrite the working token, and `DriveFolder` could give a legacy installation a new, empty folder. It is now `Bool?`; no answer stops both, and the window keeps its last answer instead of offering "Connect Google Drive". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…imit, misleading rows - One failed read of the Time Machine preferences during a backup raised "most often Full Disk Access is missing" (08.10 19:52). The watchdog now reads up to three times, 5 s apart; the window needs two failed reads in a row. Missing access fails every read, so the real alarm still comes. - The 3500 GB limit in machines.json sat under the machine key from before the folder became the key, and was looked up only under the folder - no measurement, no alarm. Both keys are read now; `set-limit` moves the old entry. - "Back up now" followed `healthy`, so an old backup greyed out the button that fixes it. It now needs only somewhere to back up to. - "Used on Google Drive" was green with no measurement, and a stale one passed for current. Amber now; no limit is neutral. - An empty queue with abandoned bands read "Everything uploaded"; the row now counts them. The queue row takes the upload verdict's tone, the free space row the card's. - docs/design.md: the measured upload volume (up to 793 GB in a rolling 24 h window), the socket, and why the Drive trash stays off. - Workflow actions pinned to commit SHAs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the findings of the 9 Oct 2026 audit: K1, W1-W3, S1-S5 and the cheap LOW items. Each fix comes with a regression test. Every test was checked against the old behaviour: I reverted each fix in turn and confirmed its tests fail (results below).
K1 (critical): rclone remote control on a private socket
What was wrong.
rclone mountlistened on127.0.0.1:5572with--rc-no-auth, and the mount runs with--drive-use-trash=false. Loopback does not keep web pages out:Origin.operations/purgeongdrive:CloudMachine/<folder>, and the delete would be permanent.The audit confirmed this against the running mount:
config/listremotesanswered a form POST sent withOrigin: https://evil.example.Fix
--rc-addr unix://~/.cloudmachine/run/rc.sock, in a directory set to 0700 (re-applied on every start).rclone rcdin /tmp: the socket works, and the client reaches it withrclone rc --unix-socket.DriveBufferService.rc(...). That coversvfs/stats,vfs/queueandvfs/queue-set-expiry, used by the GUI,drive-status,buffer-guard,backup-healthand detach. A test fails if any source file builds its own["rc", ...]call. There are no rc calls in scripts or docs.kill -9, and the next start then fails withbind: address already in use(checked on 1.75.1). Under KeepAlive the mount would never come back. Soprepare()now:prepare()stops the start with a clear message. Without that check, rclone fails withinvalid argument.--rc-no-authstays, on purpose. On a socket nothing gains from a password:vfs/queue-set-expiry, which every detach needs, requires either a password or--rc-no-auth.rclone.conf, token included, directly.--rc-addris the socket, and there is no127.0.0.1,localhostor:5572anywhere in the arguments.--unix-socketneeds rclone ≥ 1.68 (Sept 2024). Every CloudMachine install downloaded a newer one:install-rclonetakes the latest, and the project dates from 2026.Drive trash (
--drive-use-trash=true) as a second layer: not enabled.storageQuotaExceeded).operations/aboutcounts the trash.rclone purge --drive-use-trash=false.The reasoning is written up in
docs/design.mdand in a comment next to the flag.W1:
isConfiguredfails closedRemoteConfigurer.isConfiguredreturnedfalseon a timeout, a failed start or exit code 1. An unreadable or brokenrclone.confgives exit 1 (checked).Bool?, matches whole lines (mygdrive:is notgdrive:), and requiressucceeded.connectstops onnilwith a message, even with--replace-existing.DriveFolder.decide(legacyEvidence: nil)refuses whenever the answer would decide the name.Not done: the audit's extra local evidence for a legacy installation. There is no reliable local trace:
gdrive{hash};agents-versionis written at the first launch of any install.W3:
ProcessRunnerlost outputCause.
terminationHandlerremoved thereadabilityHandlers before the last data in the pipe was read. Reads also ran on a separate Foundation queue, so a chunk already read could reach the serial queue after the result had been built.Fix
DispatchSourceReadon the same serial queue that builds the result.--daemonchild that keeps the pipe open cannot hold the result.Test. 8 tasks × 500 runs of
/bin/caton a 4.4 kB file must return exactly the file every time. It also checks stderr, and that an inherited pipe (sleep 20 &) does not delay the result.origin/main: 970 of 4000 calls lost their output.TimeMachineStatus.outputnow returnsnil(no answer) when the exit code is non-zero or the output is empty. The exception is "No destinations configured", which still counts as an answer, whatever exit code comes with it. Checked here:tmutil statusandtmutil destinationinfoboth exit 0 on a working system.W2: false "no Full Disk Access" alarm
BackupHealth.readPreferences).BackupHealth.ReadConfirmation).S1-S5 and LOW
MachineBudget.limitGBlooks the limit up under the folder first, then under the stored machine key (MachineIdentity.storedKey).set-limitmoves an old entry instead of duplicating it.rclone sizeevery 6 h) and the 90%/100% alarms start working.canStartBackup: dependencies, remote, mount, attached image, registered TM,!isBusy. An old backup, an unread queue or the daily limit no longer grey it out. This also fixes N5: the menu now respectsisBusy.MachineBudget.standing: no measurement, or one older than 12 h, shows amber; no limit shows neutral. This also fixes N2, the inconsistent tone between the row and the sidebar.erroredFiles > 0shows "N fragments abandoned - only on this Mac". The queue row takes its tone fromuploadTone(uploadState)(N3), and the free-space row usesfreeSpaceTone(N2).docs/design.mdnow gives 327-793 GB per day, with the audit's measurement and its caveat about bands smaller than 32 MiB.actions/checkout,upload-artifactandaction-gh-releaseare pinned to SHAs, with the version in a comment.Skipped
Date.formattedignoringCM_LANGUAGE): needs locale plumbing fromL10ninto every formatter. Not cheap, and not risk-free.rclone.logwhile rclone runs):logShowsUploadStalledreads the last 30 minutes of that log. A copy-truncate on a running mount could cut exactly those lines and silence the daily-limit detection. It needs its own design.What happens on
brew upgradefrom 1.3.4uninstall quit:), replaces the bundle and reopens it.AgentRepair.afterLaunchreloadsbackup-health,gdrive-attachandbuffer-guard, so they run the new code.gdrive-bufferis not restarted. This is deliberate and unchanged: restarting it drops the mount under Time Machine. The running rclone (on production, pid 1758) keeps listening on127.0.0.1:5572with no password.RCTransport.legacyTCP). The queue, the watchdog and detach keep working, exactly as before the upgrade.drive-statusshows:Remote control: OPEN on 127.0.0.1:5572 - the mount was started by an older version ... run prepare-shutdown, then restart the Mac.cloudmachine-agent mount-drivefrom the new bundle. The plist does not change.prepare()creates~/.cloudmachine/run(0700), and rclone starts on the socket.cloudmachine-agent prepare-shutdown, then a restart.The other fixes take effect as soon as the app and the agents run the new code, after step 2.
Tests
swift build: OK.swift test --skip BackupImageServiceTests: 359 tests, 0 failures, locally. I skippedBackupImageServiceTestsbecause it takes the production image lock on this Mac. CI runs everything.swift format lint --strict --recursive Sources Tests: clean with the toolchain's swift-format. CI uses the Homebrew one.remoteListedback tofalseon failuredecidewithniltreated asfalseanswer(from:)ignoring the exit codehealthyProcessRunneronorigin/mainNot run: the app itself, launchd, hdiutil, a real
rclone mounton the socket. This Mac runs the production backup, so I did not touch any of them. The socket mechanism was checked on a separaterclone rcd.🤖 Generated with Claude Code