Nautilus
Run a coding agent from your phone, then review and merge its work on your PC. Self-hosted on a free cloud machine.
A phone is a good place to ask for work and a bad place to review it.
- Year
- Stack
-
- TypeScript TypeScript
- Next.js Next.js
- Node.js Node.js
- Tauri Tauri
- Git Git
The idea
A coding agent can work for twenty minutes without you. Almost every tool built around one still assumes you’re sitting at the machine it runs on.
Nautilus splits the job in two. On the phone, an installable web app lets you pick a project, send a prompt and watch the agent’s turn stream in. You approve or deny the commands it wants to run, and a preview tab shows the project’s dev server live. Then you lock the screen, and it tells you when it’s done. Your laptop can stay closed the whole time.
At your desk, a desktop app shows what changed as a real diff, file by file. Nothing reaches your files until you pull. If you edited the same file while you were away, you pick a side per file, and you can undo the last pull.
I built it in five days over the holiday. It’s open source under GPL-3.0.
Why I built it
I had two goals, and neither came first. I wanted to build something I’m proud of and would use every day, once it has everything I want in it. I also wanted practice in system design: planning something that spans several machines, making every decision myself and understanding every part of the code. Nautilus can’t be a single Next.js app, which made it a good fit for both. One machine runs a gateway, an API, the phone app, the agent and each project’s dev server. Another runs a desktop app and a sync agent.
It was also the first time I picked up technology I didn’t know and felt in control of it fast. I had never shipped a desktop app. Tauri, its permission model, Rust and building installers for three operating systems were all new to me. I kept the Rust side small on purpose. It owns the tray icon and keeps a single instance running, and every sync rule lives in TypeScript and Tauri’s capability file, where I could test it.
Working with agents
I used coding agents heavily while building it, and I don’t hide that. Most of the time they were free models, like the ones in OpenCode’s Zen collection. An agent left alone drifts, so before most of the code, I set up the project to keep it on track:
- Formatting and lint rules, plus a check for unused code
- Conventions written into the repo’s docs, so the agent reads the rules instead of guessing them
- Skills installed on my machine that every workspace can use
- Tests for each feature, and CI that runs every check before a change can count as done
- A commit convention, so the history stays readable
The planning was mine. I worked out each feature end to end before the agent touched it, including what happens when it fails halfway. The design I settled on during the first day is the one that shipped. When a result was wrong or had side effects I didn’t want, I changed the plan and worked around the problem instead of accepting it. Toward the end I did several rounds of refactoring, and I reviewed every change critically before it stayed.
Built on a $0 budget
I set one rule before writing anything. No credit card, no domain, no Docker, no paid server. The runner is a free Lightning AI Studio, a CPU machine with SSH and ports exposed over HTTPS. Free came with terms, and three of them shaped most of the design.
The machine restarts every four hours
Processes die, installed dependencies vanish, and a file written right before the restart can come back older or cut short. So Nautilus treats every boot as a full recycle. The boot script reinstalls what’s gone and takes a lock so two boots can’t race. Then the runner compares each project’s files with its last checkpoint. A project that doesn’t match is marked unhealthy and isn’t served. Showing nothing beats showing the wrong version.
A turn the agent was in the middle of comes back as interrupted, never as finished, and the phone offers a retry.
The agent never touches your Git
The agent needs checkpoints, three-way merges and undo. Your repository’s branches, hooks and remotes are yours, so they stay out of it. Each side keeps a shadow repository instead. It’s a bare Git directory whose worktree is your project folder, and your .git is never read or written.
- Sync sends a
git bundlewith only the commits since the last shared base. The receiving side checks it by SHA-256 and ancestry before importing it. - Every merge runs in full in a throwaway index first, so a conflict writes nothing.
- An apply writes a transaction record and a recovery bundle before the first file changes. If something cuts it off, the next start rolls it back.
- Every agent turn is a checkpoint. The phone can undo one turn and keep the ones after it, the way
git revertdoes.
None of this needed anything Git doesn’t already ship. The hard part was deciding what each step is allowed to assume.
The front door works against you
Lightning’s proxy rewrites every cookie to SameSite=None and adds Access-Control-Allow-Origin: * to every response. Left alone, that would let any site on the internet send requests that carry your session. So every write needs a matching Origin plus a custom header a cross-site form can’t set. Admin routes return 404 on the public side. One-time preview links only work on POST, so a link-preview bot can’t use one up.
Your laptop stays closed to the internet
The PC never accepts a connection. The desktop app dials out over SSH, and opens a reverse forward to its sync agent only during a sync you started. Tauri’s capability file lets the app start exactly three programs, each with checks on its arguments. The PC signs a grant for every sync, for one project and one direction, valid for at most 10 minutes. The runner can’t make one, and a push grant can’t write to the PC at all.
Previews still need .env values, and sync never carries .env files or keys. You pick the values in the desktop app, which keeps them in the OS keychain and sends them write-only. Only the dev server gets them. The agent runs in a bubblewrap sandbox and sees the key names only. If a value still shows up in its output, the runner replaces it with [redacted:KEY].
Where it stands
Nautilus launched on Hacker News and Product Hunt on October 4, 2026. You install it like any other open-source tool. Download the desktop app from the GitHub releases, then run one command to set up your own runner.
It has 302 tests, installers for Linux, macOS and Windows, and a monthly bill of zero.
Some of it is still rough. The builds aren’t signed, so macOS and Windows warn on first launch. Lightning is the only host I’ve tested, though anything with one HTTPS port and SSH should work. So far I’ve used Nautilus more for testing than for real work. Day-to-day use is the goal, and getting there means polishing it and adding the pieces still missing.
What I took from it
- Refuse instead of guessing. Symlinks, stale heads, timed-out requests and undos that can’t apply cleanly all get a clear no. Each refusal is a small annoyance. Each guess is a way for two machines to quietly disagree.
- Ask “what if it dies here?” at every step. The transaction record, the recovery bundle and the expiry checks all came from asking that.
- The edges are the hard part. No single service was difficult. Making sure one failing piece costs you that piece and nothing else was. When one project is unhealthy, the runner reports itself as degraded and keeps serving the rest.