# My dev setup for AI agents I hardly write code by hand anymore. Most of the time several Claude Code agents work on different repositories at once, and I read, steer and decide. For that I built myself an environment with two equally important goals: I want to **work comfortably with many agents at once**, from anywhere, from the laptop as well as the phone, without buying an expensive machine for it. And I want to do that **without handing the agents my digital identity**. These pages describe the environment. This is **not an installation guide**. Guides go stale faster than you can write them. What you'll find here is the target architecture: which building blocks exist, what role they play, what they're good for, what the alternatives are and why I chose the way I did. Where commands show up, they are illustrations. ## Video A five-and-a-half-minute animated explainer: why I rent the machine instead of buying a fat laptop, how I work with many agents in parallel from anywhere, and how the agents get by without my identity. The voice is AI-generated. ## Podcast Rather listen than read? Two hosts spend a little over a quarter of an hour talking through this setup, its architecture and the reasoning behind it. The voices are AI-generated. ## The picture ## Four load-bearing principles Almost every decision in this setup follows from one of four principles. ### 1. Rent the machine, don't upgrade the laptop The actual work is done by a rented bare-metal server in a data centre: lots of memory, around the clock, for a two-digit euro amount per month. Five to ten agents building and testing at the same time run there, not on my machine. My laptop only needs a good screen and a battery that lasts the day. The sessions live on the server: when I close the lid, the agents keep working, and I can check in from my phone or pick up later on the laptop exactly where I left off. ### 2. Agents get only as much identity as they need An agent working on repository A needs access to repository A. It doesn't need my GitHub account, my SSH keys or my access to the production server. So agents get **narrowly scoped tokens** instead of my credentials: only for the repos of their environment, and only from a socket that exists solely inside that environment. As long as it runs, the agent has continuous access through it; I can revoke that immediately at any time. ### 3. Environments are reproducible and disposable Every working environment is built from the repository's own `devcontainer.json`. It can be deleted and rebuilt at any time. Whatever has to survive (the code, the Claude login, the agent's memory) lives outside the container on the host. That's why an agent may do anything inside its environment: it *is* the sandbox. ### 4. Anything sensitive needs the human My SSH key lives in the Secure Enclave of my Mac and asks for my fingerprint every time it's used. An agent can't do anything with my identity behind my back – at most it can ask me. That's occasionally annoying and exactly the point. ## The building blocks 1. [Compute](/en/docs/dev-setup/compute): rent a powerful server instead of buying a powerful laptop. Lots of RAM for little money, and the laptop stays light, quiet and on battery for long. 2. [Network](/en/docs/dev-setup/network): a tailnet makes every environment reachable under its own name, from anywhere, from the laptop as well as the phone. It provides reachability, not security. 3. [Hatchery and drones](/en/docs/dev-setup/hatchery): one container per repository, created with a single command, and code and agent memory survive every rebuild. 4. [Identity and access](/en/docs/dev-setup/identity): two separate channels for agent and human. The heart of it. 5. [Workplace](/en/docs/dev-setup/workplace): SSH, zellij, many agents in parallel. Sessions survive dropped connections, and I keep steering from the phone. 6. [Native apps](/en/docs/dev-setup/native-apps): Android on the server, iOS as a side note. 7. [Boundary to the deploy stack](/en/docs/dev-setup/deploy-boundary): where this setup ends. 8. [FAQ](/en/docs/dev-setup/faq) All tools mentioned here that are mine are open source: [hatchery](https://github.com/levino/hatchery), [dotfiles](https://github.com/levino/dotfiles), [devcontainer-template](https://github.com/levino/devcontainer-template) and [this website](https://github.com/levino/levinkeller.de) itself. --- # Compute: rent, don't buy ## What it is A single rented bare-metal server in a data centre that runs all development environments. Mine is a second-hand machine from the Hetzner server auction: an older Intel quad-core with 62 GB of RAM. Nothing fancy, but it runs around the clock, sits on a fast line and costs a two-digit euro amount per month. **Rent, don't buy** is the most important decision in the whole setup. A laptop with 64 GB of RAM costs several thousand euros, is outdated after a few years and still has only one battery and one fan. For that money the rented server runs for years, and if it's no longer enough, I cancel it and rent a bigger one. ## Its role It's the workhorse. Every repository an agent works on gets its own container there, with its own toolchain, dependencies, dev server and tests. Five to ten of them at the same time is normal. For what agents do, **memory** is the scarce resource, not CPU time: every container holds its Node process, its language server, its test runner, sometimes a browser for end-to-end tests. 16 GB is enough to start with; with 64 GB you stop thinking about it. Many cores help when several agents build at once. ## What it buys you: rent the machine, don't buy a powerful laptop That's the real punchline. When the work happens on the server, the laptop is just a terminal with a nice screen. I work on a MacBook Neo: light, long battery life, great display, and more than fast enough for what it has to do. It doesn't get warm when ten agents run `npm install` at the same time, because that happens elsewhere. Also: with agents, I hardly type. I talk. For dictation I use [Whispering](https://github.com/epicenter-so/epicenter), an open-source speech-to-text tool. Explaining to an agent for two minutes what I want is faster and usually more precise than writing it down. And the agents keep working when the laptop lid is closed. ## Alternatives - **Locally on the laptop.** Works as long as it's one or two agents. After that the laptop gets loud, hot and drained, and every trip interrupts the work. - **Cloud VMs billed by the hour** (AWS, GCP, Hetzner Cloud). Flexible, but for a machine that runs all day anyway, much more expensive than dedicated hardware. - **Hosted environments** like GitHub Codespaces. No plain SSH access, environments get stopped when idle, billed per hour and core, and you're tied to one vendor. More on that under [Hatchery](/en/docs/dev-setup/hatchery). - **A machine at home.** Works, but depends on your home connection and your home power supply. ## Why this way A rented, used dedicated server is the cheapest way to have lots of RAM permanently available, and the most convenient: I don't have to buy, upgrade or carry any hardware. Development environments don't need high availability: if the server dies, the code is still on GitHub, and the environments can be rebuilt on any other machine from their `devcontainer.json` files. --- # Network: reachability, not security ## What it is All my devices and all development environments sit in one shared private network, a **tailnet** based on WireGuard. I use [Tailscale](https://tailscale.com) as the client. Coordination is done by [Headscale](https://github.com/juanfont/headscale), the open-source implementation of the Tailscale control server. I run it myself and log in through my own identity provider. To get started, Tailscale's free tier is just as good; Headscale is a side note for people who want to self-host everything. ## Its role Every development environment joins the tailnet as **a machine of its own**, with its own name. They all listen on the same ports and differ only by name: `hatchery-levino-shipyard`, `hatchery-levino-levinkeller-de` and so on. Instead of remembering which environment is on which port of the server, I address it directly. I don't hand out ports on the server, don't maintain port forwards and don't open anything to the internet. To look at an environment's dev server in the browser, I just open its name plus the port, from the laptop or the phone. The environments are registered as **ephemeral** nodes: delete one, and after a while it disappears from the device list by itself. ## What it is *not* The tailnet is **not** my security layer. It's convenient and it reduces the attack surface, but I don't rely on it. Being in the tailnet doesn't get anyone into an environment: each one requires my key for SSH login, and that key needs my fingerprint (see [Identity and access](/en/docs/dev-setup/identity)). Password login is off. There's a practical reason for that: a network that gets misconfigured once, a device that gets lost, a share you forgot about – all of that happens. If security depends on a single layer, that's one too few. ## A back door, on purpose The server itself is also reachable via plain SSH from the internet. That's deliberate: when the tailnet acts up (and it occasionally does), I can still get in and fix it. A network that can only be repaired through itself is a trap. ## Alternatives - **Opening ports on the host** and working through SSH tunnels. Works, but scales poorly: every environment needs its own ports and you have to remember them. - **A classic VPN** (OpenVPN, plain WireGuard). Works, but you maintain keys and IP addresses by hand, and new environments don't show up on their own. - **ZeroTier, Nebula, NetBird.** Similar idea to Tailscale. Tailscale won for me because there's a ready-made devcontainer feature and because Headscale exists as a free control server. --- # Hatchery and drones ## What it is [Hatchery](https://github.com/levino/hatchery) is a small open-source tool I wrote for exactly this setup. It manages development environments on the server: one devcontainer per repository (or per task). In Hatchery speak these containers are called **drones**. The whole tool is named in StarCraft Zerg style because that's fun: drones get spawned, burrowed (`burrow`), unburrowed (`unburrow`) and slain (`slay`). ## How a drone comes to life ```bash hatchery spawn levino/shipyard ``` Hatchery clones the repository onto the host and starts a container from it with the official [devcontainer CLI](https://containers.dev), using the `devcontainer.json` that lives **in the repository itself**. On top, it sneaks in a few features: - an SSH server so I can connect, - Tailscale so the drone becomes [a machine of its own in the tailnet](/en/docs/dev-setup/network), - the GitHub CLI, - a Hatchery feature of its own: Claude Code, [zellij](https://zellij.dev), the credential helpers for Git and `gh` (see [Identity and access](/en/docs/dev-setup/identity)) and a fallback `CLAUDE.md` with ground rules for agents. The repository needs **nothing Hatchery-specific** for this. The same `devcontainer.json` works unchanged in GitHub Codespaces or locally in VS Code. Hatchery is replaceable; the repos stay portable. Optionally Hatchery also loads my [dotfiles](https://github.com/levino/dotfiles) into every drone: shell configuration, Git settings, a global `CLAUDE.md` and skills with my coding conventions. ## What survives and what doesn't A drone is disposable. Whatever must not be disposable lives on the host and is mounted in: - **the code**, as a Git worktree on the host, - **Claude Code's state**: login, settings, memory and history. So you can rebuild a drone from scratch (say, because the `devcontainer.json` changed) and the agent still knows what it was working on. I only have to log in to Claude once per new drone. ## Docker is the single source of truth Hatchery has no database of its own. Which drones exist is stored in the labels of the Docker containers and nowhere else. `hatchery list` asks Docker. So Hatchery's state can never drift from reality, and you can handle drones with plain Docker commands too. Besides the CLI there's one small service that runs permanently: the **credential service**. It watches Docker events and sets up GitHub access for every drone that starts. That's the interesting part and it has [a page of its own](/en/docs/dev-setup/identity). ## Alternatives and why not Before writing Hatchery I tried what exists: - **GitHub Codespaces.** What killed it was the way of working, not the price. I can't just SSH into a codespace like into any other machine. There's only `gh codespace ssh`, a tunnel through the GitHub CLI that also needs an SSH server in the image. So I ended up running Claude Code in the VS Code terminal, and that was painful: rendering glitches, sluggishness, and again and again the connection dropped and the session was gone. On top of that, Codespaces stops an environment once it sits unused for a while (default: 30 minutes idle). Agents working in the background for hours don't fit that model. Then there's the money: billed by hours and cores, fixed machine sizes, tied to GitHub. Ten environments running in parallel all day gets expensive. A drone, by contrast, is a perfectly normal machine on the tailnet: `ssh` in, start zellij, and the session survives every dropped connection. Nothing gets stopped. - **DevPod.** Good idea, but the state lives on the client: the laptop you created the environments on is the one that knows about them. I want to see the same thing from the phone, the laptop and the server. - **Coder and similar platforms.** Powerful, but they bring Kubernetes or a platform of their own that you have to operate. Too much for one person with one server. - **Plain Docker by hand.** Works, but then you keep rebuilding by hand exactly what Hatchery automates: joining the tailnet, SSH keys, tokens. And one thing spoke against all of them: none of the alternatives has a good answer to how an agent in the environment gets to GitHub without getting my full identity. --- # Identity and access This is the heart of the setup. Everything else is convenience; here it's about one question: **What may an agent that has full rights inside a drone do outside that drone?** My answer: there are two separate channels. The agent works through one, I work through the other. ## The agent channel: a GitHub App instead of my identity ### The problem The obvious way is to run `gh auth login` in the environment or to pass in your own SSH key. Then the agent can push. But it can also do everything else I can do: write to every one of my repositories, in every organisation I'm a member of, delete releases, change settings. An agent that goes off the rails, a booby-trapped README or a compromised dependency – and the damage isn't limited to the repository being worked on. Fine-grained personal access tokens would be an alternative, but creating, expiring and renewing them by hand per environment is tedious and error-prone. ### The solution I created my own **GitHub App** and installed it in my organisations. A GitHub App can issue **installation tokens** that can be limited to single repositories. They don't belong to my user account but to the app. Only the credential service on the host knows the app's private key. It mounts a **Unix socket** into every drone. Whoever asks at that socket gets a fresh token, but only for the repositories granted to exactly this drone. If the drone asks for another repo, it gets a `403` along with a hint on how I could grant it. There are no passwords and no tokens on the drone's disk. **The identity is the mount itself**: which drone is asking follows from which socket it is. That can't be faked, because the host creates the socket, not the container. So as long as the drone runs, the agent has continuous access to its repos: every access fetches a fresh token, and none is ever stored. That installation tokens expire after one hour isn't the boundary, just a safety net in case one does get out of the drone: then it's useless after an hour at most and only ever worked for those repos. A `gh auth login` token or my personal SSH key, by contrast, works everywhere and doesn't expire. ### How Git and `gh` find out The Hatchery feature sets up two things in every drone: - a **Git credential helper** that asks the socket for a token on every GitHub access. SSH URLs are rewritten to HTTPS so `git@github.com:…` works too. - a **wrapper around `gh`** that fetches a token before every call. To the agent it feels like a normally logged-in system. `git push` and `gh pr create` just work. Every drone also gets a `CLAUDE.md` with three rules: never `gh auth login`, never hardcode tokens, and on authentication errors say so instead of working around them. ### Grants change at runtime If an agent needs a second repository, I grant it: ```bash hatchery repo connect levino/shipyard levino/levinkeller.de ``` That takes effect immediately, without a restart. Taking it back is just as quick: `hatchery repo disconnect` removes a repo from the grants, `hatchery slay` removes the drone along with its socket. I don't have to wait for any token to expire. The list of grants lives on the host **outside** the drone. So a drone can't widen its own rights, not even by editing a file and waiting for the next restart. ### A stricter model for comparison For repositories on my own [Forgejo](https://forgejo.org) instance, Hatchery goes one step further: the drone doesn't get a real token at all, only a placeholder. Its Git traffic goes through a per-drone proxy that checks every request against the grant list and only then swaps in the real token. The agent never sees the secret. That's cleaner, but also more effort – for GitHub the token model is enough for me. ## The human channel: SSH with a fingerprint ### My key never leaves the Mac My SSH key lives in the **Secure Enclave** of my Mac, managed with [Secretive](https://github.com/maxgoedjen/secretive). It can't be exported, copied or read out. Every signature with it asks for confirmation via Touch ID. I use this key to log in to the drones. When they start, the drones fetch the allowed public keys straight from my GitHub profile. ### Agent forwarding on a leash I often connect with `ssh -A`, i.e. with agent forwarding. That lets the drone *use* my key, for example to push to a repository that doesn't go through the app, or to hop on to another server. Classically that would be dangerous: any process in the drone could log in anywhere as me for as long as the connection is open. With Secretive it's different. **Every single use** of the key pops up on my Mac and wants my finger. So an agent can't quietly hop onto the production server. At most it can ask me, and then I see what it's up to. To be honest: sometimes I loosen the leash. When an agent is supposed to work on a server for an hour, I keep a shared SSH connection (`ControlMaster`) open so I don't have to confirm every minute. That's a conscious decision for a limited time, not a permanent state. ## Why two channels The two channels cover different needs: | | Agent channel | Human channel | |---|---|---| | **Who** | the agent, any time | me, with my finger | | **For** | reading and pushing code, PRs, issues | logging in to drones, servers, anything sensitive | | **Reach** | only granted repos | everything I'm allowed to | | **If it leaks** | token works for 1 hour at most, only for these repos | key never leaves the Mac, signing only with my finger | | **Revoke** | immediately via `repo disconnect` or `slay` | don't put my finger down | | **Without me** | runs | stops | The agent can do its work without me. Anything beyond that needs me physically. That difference is the core of the whole setup. --- # Workplace ## At the laptop: SSH, zellij, several agents My workplace is a terminal. I SSH into a drone and start [zellij](https://zellij.dev) there, a terminal multiplexer (like tmux, but with friendlier defaults). zellij runs several tabs and panes, and in several of them a Claude Code agent is at work. The agents spin up sub-agents of their own for research or reviews when needed. zellij has an important side effect: the session lives on in the drone when the connection drops. Close the laptop, the train enters a tunnel, whatever – next time I connect everything is still there, and the agents kept working in the meantime. That's exactly what I was missing with GitHub Codespaces: there Claude Code ran in the VS Code terminal, and when the connection dropped, the session was gone. Usually I have several drones open at once, one per project. I switch between them the way you switch between colleagues: who's done, who has a question, who needs a decision? ## Agents without confirmation prompts By default Claude Code asks for permission before every command and every file change. In the drone I turn that off (`--dangerously-skip-permissions`). That sounds more dangerous than it is, because **the drone is the sandbox**: - It's disposable and can be rebuilt from the `devcontainer.json` at any time. - The code lives in Git, and everything important lands on GitHub through pull requests. - Its GitHub token only reaches the granted repositories. - Anything that needs my identity needs my finger. An agent that keeps asking for permission isn't an agent, it's a very slow colleague. Safety comes from the environment, not from the prompts. ## What agents bring along Every repository has a detailed `CLAUDE.md` with project knowledge, often plus skills and sub-agents of its own in `.claude/`. My personal conventions (for example test-driven development, functional style with [Effect](https://effect.website), no classes) live as a skill in my [dotfiles](https://github.com/levino/dotfiles) and reach every drone via Hatchery. ## On the go: Remote Control instead of SSH From the phone I do **not** connect via SSH. A terminal on the phone is a pain, and my key lives on the Mac anyway. Instead I use Claude Code's **Remote Control**: I switch it on in a running session, and then I can continue that session in the Claude app on my phone. I see what the agent is doing, answer its questions, give it the next task. That works well for what you actually do on the go: read, decide, say something short. Voice input on the phone fits right in. The catch: when the agent needs my SSH key, I have to confirm on the Mac. That doesn't work from the phone. In practice it rarely matters, because the agent's work runs through the [agent channel](/en/docs/dev-setup/identity) and doesn't need me. ## Alternatives - **VS Code Remote-SSH or JetBrains Gateway.** Works with every drone because each has a perfectly normal SSH server. I rarely need the IDE because I hardly edit code myself. - **tmux instead of zellij.** Just as good if you know it. - **Mobile SSH clients** like Blink or Termius. They work, but an agent wants to be read and steered, not operated. A chat interface is better for that. --- # Native apps For web projects a container is enough. Mobile apps need emulators, and emulators need virtualisation. ## Android: on the server Android emulators run on Linux but need hardware virtualisation (KVM) to be usably fast. Hatchery can pass the `/dev/kvm` device into a drone: ```bash hatchery spawn levino/some-app --kvm ``` The emulator then runs right inside the drone, and the agent can build, install and test the app like any other process. This is deliberately **opt-in**, because almost no drone needs it. ## iOS: level 2 iOS apps can only be built and tested in the simulator on macOS – Apple's rule. For that I run a macOS VM on my Mac Studio with Xcode, the simulator and `xcodebuild`, which an agent reaches via SSH. It has its own narrowly scoped access, just like a drone. It's hand-built and not automated by Hatchery. macOS VMs can't do nested virtualisation, so there's no Docker inside either. For this setup it's a side note: it shows that the principle (own environment, own identity, agent may do anything inside) carries beyond Linux containers. --- # Boundary to the deploy stack This setup ends where code reaches GitHub. How it gets from there to production is a system of its own: Kubernetes (k3s), GitOps with Argo CD, identities via ZITADEL, everything described as code. That will be a separate part of these docs. The two systems are separate but touch in exactly two places. ## Touch point 1: Git and CI Agents push branches and open pull requests. From there CI takes over: it builds, tests and creates preview environments. This website, for instance, gets its own preview at `pr-.levinkeller.de` for every pull request. After the merge, CI builds an image and records the new version on a deploy branch that Argo CD rolls out in the cluster. An agent needs **no access to the cluster** for any of this. It only needs its repository. That's the main reason the token model from [Identity and access](/en/docs/dev-setup/identity) is enough. ## Touch point 2: my SSH key If someone really has to get onto a production server directly, to look at or fix something, that only works through my key. And therefore through my finger. Agents can ask me to do it; they can't do it themselves. ## Outlook Part two will describe the deploy stack: how services come to life, how they get identities and tokens, and how to set up a new stack from a template (a [Copier](https://copier.readthedocs.io) template called `agentops-community-stack`). --- # FAQ ## Why not just `gh auth login` in the environment? Because then the agent *is* me. The token from `gh auth login` is valid for every repository in every organisation I have access to, and it doesn't expire. Instead, the drone fetches a fresh GitHub App token on every access, valid only for its repositories and useless after an hour at most should it ever get out. More under [Identity and access](/en/docs/dev-setup/identity). ## Why no SSH keys in the drones? Same reason. An SSH key for GitHub is bound to my account and opens everything. A key lying around in a drone can also be copied. My only key lives in the Secure Enclave of my Mac and can't leave it. ## But you forward the SSH agent with `ssh -A`? Yes. That's the human channel. The drone can use my key, but only with my fingerprint, for every single signature. An agent can't do anything with it behind my back. Sometimes I deliberately keep a connection open for a while so I don't have to confirm constantly – that's a decision for a limited time. ## Isn't the tailnet the actual security layer? No. The tailnet makes every drone reachable under its name. Anyone in the tailnet still needs my key to log in. And the dev server is deliberately reachable via SSH without the tailnet too, as an emergency exit when the tailnet acts up. See [Network](/en/docs/dev-setup/network). ## Why not GitHub Codespaces or DevPod? With Codespaces the way of working was the problem. I can't SSH into one like into any other machine, only through the `gh codespace ssh` tunnel. So Claude Code ran in the VS Code terminal, with rendering glitches, sluggishness and dropped sessions. And after an idle timeout (default: 30 minutes) Codespaces stops the environment, which doesn't fit agents working in the background for hours. It's also expensive for many always-on environments and tied to GitHub. A drone I reach with plain `ssh` over the tailnet, the zellij session survives dropped connections, and nothing idles out. DevPod keeps the state on the client, and I want to see the same thing from every device. And neither has a good answer to how the agent gets to GitHub without getting my identity. The repositories stay compatible though: every `devcontainer.json` Hatchery uses also works in Codespaces. ## Isn't `--dangerously-skip-permissions` dangerous? On your own laptop: yes. In a drone: hardly. The drone is disposable, its token only reaches its repositories, and anything involving my identity needs my finger. The worst an agent can do is a broken branch in a granted repository. ## Why a weak laptop? Because it doesn't have to compute anything. The work happens on the server. What I need from the laptop is a good display, a good keyboard, a good microphone and long battery life. See [Compute](/en/docs/dev-setup/compute). ## Do you type all of that? No, I talk. With speech-to-text I explain to an agent in two minutes what I want. That's faster than typing and usually more precise too, because you give more context when you speak. ## And from the phone? No SSH. I switch on Remote Control in a running Claude Code session and steer it from the Claude app. Only when the agent needs my SSH key do I have to go to the Mac. See [Workplace](/en/docs/dev-setup/workplace). ## What about iOS apps? They need macOS. I have a macOS VM on a Mac Studio for that, which agents reach via SSH. It's not automated and more of a side note, see [Native apps](/en/docs/dev-setup/native-apps). Android, on the other hand, runs right inside a drone on the server. ## What does it cost? The server is a used dedicated machine from the Hetzner server auction and costs a two-digit euro amount per month. Tailscale is free for personal use, Headscale and Hatchery are open source. The biggest item is the subscription for the agents themselves. ## Can I use Hatchery? Yes, it's [open source](https://github.com/levino/hatchery). It is, however, built for exactly my use case: one human, one server, GitHub. But the ideas carry over without Hatchery too: a GitHub App for agent tokens, a hardware-bound SSH key for the human, and environments you can throw away.