CI runners — cloud, local, and the switch between them¶
Outcrop's CI can run on GitHub-hosted runners or on a self-hosted Mac. This page covers how the switch works, how to set the Mac up, and the rule for deciding which work belongs where.
Why this exists¶
GitHub-hosted minutes are metered; a self-hosted runner is not. Billing
applies to GitHub-hosted runners only, so a self-hosted runner keeps
working after the plan's minutes are exhausted — which is exactly the
situation that motivated this page (see RUNLOG, 2026-07-31: every job on
every branch failing in ~3 seconds with runner_id: 0 and no logs, because
no runner was ever assigned).
Two failure modes look identical from the outside and are worth separating:
- Minutes exhausted. GitHub-hosted jobs fail instantly; self-hosted jobs still run. Flipping to the Mac restores CI.
- Actions disabled for the account (billing failure, org policy). Nothing runs anywhere, and a self-hosted runner will not save you.
Setting up the Mac runner and seeing a green job is what tells you which one you are in.
The switch¶
runs-on is a repository variable, not a hardcoded label:
runs-on: ${{ fromJSON(vars.RUNNER_LINUX || '"ubuntu-latest"') }}
Unset — the normal state — means GitHub-hosted. Setting the variable moves every job that reads it, with no commit and no workflow edit:
| Variable | Unset (default) | Fallback value |
|---|---|---|
RUNNER_LINUX |
ubuntu-latest |
["self-hosted","outcrop-mac"] |
RUNNER_MACOS |
macos-latest |
["self-hosted","outcrop-mac"] |
The value must be valid JSON: a quoted string for one label, or an array for several. An invalid value fails the workflow at parse time, before any job starts.
# fall back to the Mac
gh variable set RUNNER_LINUX --body '["self-hosted","outcrop-mac"]'
gh variable set RUNNER_MACOS --body '["self-hosted","outcrop-mac"]'
# back to cloud when minutes reset
gh variable delete RUNNER_LINUX
gh variable delete RUNNER_MACOS
Or by hand: Settings → Secrets and variables → Actions → Variables.
claude.yml deliberately does not read these variables. Every other
workflow runs a fixed, reviewed script; that one runs an agent that decides
its own commands from comment text. It stays on a disposable cloud VM.
Setting up the Mac¶
1. Check the toolchains first. The runner will happily accept jobs it cannot complete:
./scripts/runner-doctor.sh
It separates two things that look alike in a terminal but are not:
FAIL— nothing in the workflow installs this, so the job goes red. These are the only lines that matter, and they set the exit code.note— the workflow provisions it itself (dtolnay/rust-toolchainfor rustfmt/clippy/llvm-tools,taiki-e/install-actionfor cargo-llvm-cov,actions/setup-*for java/node/dotnet/python, the guarded brew steps for protoc/xcodegen). Having it locally only makes the first run faster.
On a Mac used for Outcrop development there is usually exactly one blocking gap, because GitHub's Ubuntu image ships an Android SDK and macOS does not. The lightweight fix — no Android Studio needed:
brew install --cask android-commandlinetools
export ANDROID_HOME=/opt/homebrew/share/android-commandlinetools
sdkmanager --licenses
Do not stop at your shell profile. The runner is a launchd service; it
never sources .zshrc or .zshenv, so an export that works perfectly in
your terminal is invisible to every job. Runner environment goes in a .env
file in the runner's own directory, and the service must be restarted to pick
it up:
cd ~/actions-runner
echo 'ANDROID_HOME=/opt/homebrew/share/android-commandlinetools' >> .env
echo 'ANDROID_SDK_ROOT=/opt/homebrew/share/android-commandlinetools' >> .env
./svc.sh stop && ./svc.sh start
This is the general shape of self-hosted debugging: "it works in my shell"
proves nothing about the runner, because they do not share an environment.
If a job cannot find a tool you know is installed, check .env and the
runner's .path file before suspecting the workflow.
Installing llvm-tools-preview and cargo-llvm-cov locally is worth it too,
even though the coverage job would install them itself — it saves that
download on every run:
rustup component add llvm-tools-preview && cargo install cargo-llvm-cov
2. Register the runner. In the repo: Settings → Actions → Runners →
New self-hosted runner → macOS → ARM64. That page generates a
registration token (valid one hour) and the download commands. Then, from
the actions-runner directory it tells you to create:
./config.sh --url https://github.com/jeffreyclegg/outcrop \
--token <TOKEN-FROM-THAT-PAGE> \
--name outcrop-mac \
--labels outcrop-mac \
--unattended --replace
--labels adds to the automatic ones (self-hosted, macOS, ARM64); the
custom outcrop-mac label is what the variables above target, so a second
machine can join later without a workflow change.
3. Run it as a service so it survives logout and reboot:
./svc.sh install
./svc.sh start
./svc.sh status
4. Keep the machine awake. A sleeping runner does not fail jobs — it leaves them queued, which reads as a hung CI:
sudo pmset -a sleep 0
5. Verify with a cheap job before trusting it:
gh variable set RUNNER_LINUX --body '["self-hosted","outcrop-mac"]'
gh workflow run docs.yml # or push any commit
The run should show outcrop-mac as the runner name.
Things that differ from a cloud runner¶
- The workspace persists.
actions/checkoutcleans the repo, buttarget/,.build/, and Homebrew state survive between runs. That is the main speed win, and the main source of "works in CI, fails on a clean machine" drift. The Apple workflow'sbrew installsteps are guarded withcommand -vso they do not pay a round trip every run. - Rust CI runs on macOS ARM for the first time. The workspace suite has
only ever run on Linux in CI. Failures there may be real macOS bugs rather
than infrastructure — exactly like the
/private/varcanonicalization bug PR #7 fixed. Treat a new failure as a finding, not a runner problem, until you have read it. - One machine is a serial queue. Jobs that fan out in parallel on cloud runners will run one after another. Register additional runners on the same Mac (separate directories, same labels) if that becomes the bottleneck.
- Disk grows. The runner's
_work/keeps a checkout and build tree per repository. Prune it when disk gets tight; it is all reproducible.
Security boundary¶
Self-hosted runners execute whatever the workflow says, on your machine, with your files. That is acceptable here because this repository is private, so only trusted pushes trigger it.
If Outcrop is ever made public — including for the free-minutes reason — this changes immediately: a pull request from a fork can propose a workflow that runs arbitrary code on your Mac. Before flipping visibility, either remove the self-hosted labels or gate every self-hosted job on the PR coming from the repository itself:
if: github.event.pull_request.head.repo.full_name == github.repository
The ruleset: what runs where¶
The decision is not "cloud or local" globally. Four properties decide it per job:
1. Cost multiplier. Linux bills 1x, Windows 2x, macOS 10x. The Apple jobs are the entire reason the minutes ran out, and they are the jobs a local Mac runs better — free, and faster on warm caches.
2. Fidelity. A job should run on the platform whose failures you care about. The Linux jobs are cheap and their platform is the one CI has always used; the Mac is a fallback for them, with the side benefit of catching macOS-specific bugs. The Apple jobs only ever meant anything on macOS.
3. Availability. Cloud runners are always there. A laptop sleeps, travels, and runs out of battery. Anything that must run unattended — the weekly OKF spec watch, anything that gates a merge while you are away — is safer in the cloud, because a self-hosted miss is a silent queue rather than a red X.
4. Trust. Fixed reviewed scripts are fine locally. Agent-driven workflows and anything handling credentials you would not want on a personal machine belong on disposable infrastructure.
Applied to what exists today:
| Job | Default | Why |
|---|---|---|
clients-apple.yml (swift build, XCUITest) |
Mac, permanently — once proven | 10x billing, and the local machine is faster and more faithful |
ci.yml (fmt, clippy, test, coverage) |
Cloud, Mac on fallback | Cheap at 1x, and Linux is the platform CI has always asserted |
docs.yml, clients.yml |
Cloud, Mac on fallback | Cheap, and the Android job needs an SDK the Mac lacks |
okf-spec-watch.yml (weekly cron) |
Cloud | Must run unattended; a sleeping Mac silently misses it |
claude.yml |
Cloud, never local | Agent-authored commands; needs a disposable VM |
Two rules worth stating outright, because they are the ones that bite:
- Prefer the cloud for anything whose absence is invisible. A failed job shouts; a queued job on a sleeping laptop does not.
- Move a job to the Mac permanently only when local is better on cost and fidelity — which today means the Apple jobs and nothing else. The rest is fallback, and the fallback should be temporary by default: flip the variables back when minutes reset, so drift between "CI passes" and "CI passes on a clean machine" has less time to accumulate.