labs · internal infrastructure
The Software
Factory.
I file an issue. It ships a pull request.
A GitHub issue goes in. A reviewed, gated, CI-green pull request comes out, opened by an agent with its own identity, on a server in my house. My part is the bit that is left: open the pull request, read the diff, decide whether it ships. Everything after that merge is automatic. It built a working app in a weekend that way. It also spent fifteen dollars once producing absolutely nothing, and the interesting part of this page is why.
98 runs between 11 and 29 July 2026. 72 shipped a pull request. 10 died as silent no-ops, 4 ran out of turns, and 12 early runs predate the verification guard, so they cannot honestly be counted either way. That is 84% of everything the guard could actually check. All of it read straight out of the job ledger, recorded by the runner's own telemetry rather than self-reported by the agent. The $270 is notional model spend: the real cost is one monthly Claude subscription that the factory shares with everything else I use it for. Not free, just already paid for. Snapshot as of 29 July 2026.
01 · why
I lead product marketing for Upsun Dispatch. You cannot market a thing you have not felt, so I went and built my own version of the problem it solves, at home, and got it wrong in public for nineteen days.
That is the honest reason this exists. The other reason is that it was genuinely fun. This is a nights-and-weekends project, not a product, not a side business, and not something I am going to help you set up. It has exactly one user.
What it does. A labeled GitHub issue triggers a self-hosted runner. A manager agent scopes the work, reads the code itself, delegates to a small crew, runs the gates, and opens a pull request as its own GitHub identity. CI runs on the PR. Then it stops and waits for me to review it. Once I merge, the deploy runs itself.
What broke. Everything on this page that works was preceded by something that did not. A manager that waited fifteen turns for a phone call that could never arrive. A dashboard frozen for seventeen hours behind a green status light. A guard I built to catch lies that then lied in the other direction. Those failures are the part that transferred to the day job, so they get most of the room below.
The definition of “it worked” does not live in the model's mouth. It lives in a script that asks GitHub what actually shipped.
The rule the whole thing rests on
Standing on
- Claude Codeheadless, on a subscription. The thing that makes the economics work at all
- Telegramphone-first intake, through terranc's open-source Claude Code bridge, which already did the hard part
- GitHub Actionsself-hosted runners, one per repo
- GAALconfig sync, so the agent's brain lives in git. One of my own labs, which meant every rough edge became a real issue on my own tracker
- Playwrightthe visual gate
- SQLite + Datasettethe job ledger and the dashboard
- Proxmoxthe home server it all runs on
Almost none of this is mine. The two best decisions in the whole build were both decisions to write less code: drop the heavy agent runtime the original plan assumed, and use an existing open-source Telegram bridge instead of hand-rolling a polling loop. Both came from stopping to ask whether I actually needed the thing I was about to build, and the answer was no twice. The interesting work was never writing components. It was wiring existing ones together and then finding out where the wiring lies to you.
This page was built by the factory. The plan became GitHub issues, the issues became pull requests, and I reviewed and merged them.
02 · the loop
Eight stages. Exactly one of them is a human, and it is deliberately the last one that can change anything.
- 01 Intake
- I describe a job from my phone. A Telegram bot turns it into a properly scoped GitHub issue, which matters more than it sounds: a badly scoped issue is the single most expensive failure mode this system has.
- 02 Trigger
- I add one label. Nothing runs without it, and labelling several repos at once runs them at once.
- 03 Runner
- A self-hosted runner picks the job up on a home server, isolated from the rest of the network, so agent-authored code never touches anything that matters.
- 04 Build
- A manager agent on Opus reads the code itself, then delegates implementation and adversarial review to a crew on Sonnet and Haiku. It is allowed to decide a subagent is not worth the handoff, and it regularly does. The models are named by alias rather than pinned version, so when Opus 5 shipped the factory picked it up on the next run with no config change at all: the ledger shows 76 runs on 4.8 and then 22 on 5, with nothing in between but a date.
- 05 Gates
- Lint, typecheck and build run before the push. In CI, a Playwright job drives every key route at desktop and mobile, screenshots them, and asserts the things a diff cannot show you.
- 06 Pull request
- The agent opens the PR as itself, through its own GitHub App, on a branch named after the issue. It is structurally incapable of merging its own work.
- 07 Review
- Me, in GitHub. I open the pull request, read the diff, check the gates went green and merge it if I am happy. Usually a couple of minutes. Sometimes I send it back. This is the only place a human can change the outcome, which is exactly why it is the last step rather than an earlier one.
- 08 Deploy
- Nobody. Merging to main builds the image, backs up the database, pulls the new digest, migrates on boot, gates on a health check, and rolls back automatically if that check fails.
Why Telegram, though
Strictly speaking I do not need it. Any Claude session with GitHub access can file and tag an issue, and sometimes that is exactly what I do. But my ideas do not arrive at a desk. They arrive on a walk, on the loo, or in the shower, which a waterproof phone turns into a legitimate place of work. The bot is the difference between an idea that becomes a pull request and one that evaporates on the way back to the office.
Wide drawing. Swipe or scroll it sideways on a small screen.
03 · trust
The agent was the easy bit. It wrote its first real feature in under a minute. Everything after that was deciding what happens when it is wrong.
Every run exits through exactly one of five verdicts, and which one it gets is decided by a single question the agent does not get a vote on: is there a pushed branch.
- shipped Branch pushed
The run was clean and the pull request is open.
- success-degraded Branch pushed
It shipped, but burned its turn budget getting there. Still a PR, still honest about what it cost. The yellow is deliberate.
- needs-input Nothing pushed
Hit a real wall, refused to guess, tagged me with the diff and the options.
- agent-failed Nothing pushed
A loud failure. Label flipped, comment posted, notification sent.
- silent-noop Nothing pushed
Claimed success and produced nothing. The dangerous one.
Every reliability fix here has been the same move. Stop trusting what the thing says about itself, and go look at what it actually produced. That is not an AI lesson. It is engineering, arriving in a new costume.
04 · observability
The first version had no windows. It worked, and I had no idea what it was doing, which is functionally the same as it not working.
Repos
- gqwebsite
- private-app
- kinhold-landing
- qthirtytwo
- private-site
- q32-factory
Running now
- 00:04reading runbook and issue #41
- 00:22scouting src/components
- 01:10implementer editing the footer
- 01:48reviewer running an advisory pass
The live status comes from the runner itself, not from polling an API. The runner already knows when a job starts and what it is doing, so a start hook and the manager's own narration write that to a local database, which is copied out over the management path. The isolation stays intact. Nothing had to open a hole to make the machine observable.
It paid for itself on the first run. The feature built to watch the factory work is exactly what caught it not working: the stream showed a manager delegating and scouting, and the ledger showed it quietly producing nothing.
05 · build log
Six moments that changed the design, in the order they happened.
- 11 July
One issue, one pull request.
Provisioned the runner and wired auth so the agent draws on a subscription rather than a metered API key. The first job failed immediately on a missing config line. The retry produced a clean branch in thirty-six seconds.
The first failure is never the scary one. - 11 July
Isolate the blast radius before widening the queue.
The runner executes agent-authored code on the same machine as a live production site and the family network. I locked its outbound access down from inside the container rather than at the host, because the host-level fix could have taken down the very thing it was protecting.
Pick the mitigation that cannot break what you are guarding. - 12 July
A stray thought becomes a shipped fix.
The site described one of my own projects wrong. Instead of opening an editor I filed it as a job. The agent found the stale entry, made the edit, opened a pull request, and I merged it. Live in minutes.
This is the entire thesis, tested on something that actually mattered. - 13 July
Killing the silent no-op.
A run that claims success while shipping nothing is the worst bug a can-I-trust-this system can have, because it is invisible and it writes its own false record. Fixed in three layers, the load-bearing one being a script that asks GitHub whether a branch exists and fails the run when it does not.
Ground truth is the artifact, not the report. - 18–19 July
A whole app in a weekend.
Pointed it at a greenfield project with its own spec and backlog. 54 runs and $174 later it had shipped the scaffold, the tooling, both CI gates, a Dockerfile, the auth schema, the feature routes and a deploy agent, each as its own reviewed pull request. My entire contribution was reading diffs and deciding which ones shipped. Nine of those 54 runs were silent no-ops, which is how I found the worst bug in the system.
The agent is the easy part now. The factory is the hard part. - 27 July
It starts building itself.
Onboarded the factory’s own repository, behind a CI check that fails any agent pull request touching the workflows, the brain or the guards. It shipped a dashboard feature and never went near the fence.
A factory that builds itself is a claim that the safety model is real rather than decorative.
06 · what broke
Each one is the same bug wearing a different hat: something trusted a proxy for the truth instead of the truth.
| Run | Cost | What it bought |
|---|---|---|
| Cheapest shipped change | $0.39 · 12 turns | 70 seconds from label to open pull request. This is the one that makes filing a job feel free, which is its own kind of trap. |
| Most expensive failure | $7.71 · 51 turns | Ran out of turns and shipped nothing. Then I retried the same issue and it did it again for $7.63, so one ticket cost $15.34 and produced no code at all. |
| Shipped and failed | $6.83 · 51 turns | Opened a clean, mergeable PR and then burned its remaining budget trying to prove to itself that the gate would pass. The only run ever marked success-degraded. |
| Ten silent no-ops | $11.84 combined | Nine of them on two consecutive days. Cheap, quiet, and reporting success, which is exactly why they were the most dangerous thing this system has ever done. |
| The heaviest day | $104.98 · 31 runs | 19 July, the second day of building an app from scratch. More than a third of the entire bill, in one sitting. |
| Average run | $2.81 · 31.3 turns | Across all 98. Notional: the real line item is one monthly subscription, not 98 invoices. |
Twenty-six of those ninety-eight runs never shipped anything. That number is here on purpose, and it is the one I would have been most tempted to round. A system that claims a hundred percent is a system that has not yet been caught lying, and I know that because mine did, ten times, cheaply, while reporting success.
07 · what it was for
The factory was the excuse. The point was learning what agentic delivery feels like from the inside, in a place where the stakes were a personal website rather than someone's production estate.
Three things came out of it that I could not have got from a demo or a briefing document. That the hard problem is not code generation, it is knowing whether the thing in front of you actually happened. That the most expensive failures are scoping failures, filed by a human, before any model runs. And that trust in a system like this is built almost entirely out of what it does when it is wrong, not what it does when it is right. All three now show up in how I talk about Upsun Dispatch, which is the job this was really for.
- More repos
- Every side project onto the same rails, so a fix is a sentence typed from my phone.
- Harder gates
- Visual acceptance criteria compiled into real assertions, so a claimed element cannot ship invisible.
- Better intake
- The most expensive failures have all been scoping failures, never coding ones. That is where the next win is.
- Write it up
- The long-form version of this log, because most of what gets written about agent reliability is written by people who have not yet had one lie to them.
08 · the obvious question
Two answers, and which one you get depends entirely on why you are asking.
If you want to understand it
Build a small one. Point it at something you own and do not much care about, and let it fail in front of you. You will learn more in a fortnight than from any amount of reading, and most of what you learn will be about plumbing rather than about models.
That was the whole point for me, and I would do it again tomorrow.
If you are a company shipping software
Genuinely, no. Look at the stack list near the top of this page, then count the things that have to keep working: runners, tokens, a control repo, a sync daemon, per-repo secrets, an isolation policy, a ledger, a dashboard, two CI gates, a guard script and a bot. Every one of them has failed at least once.
I maintain all of that for an audience of one, in the evenings, and it still lied to me ten times. Multiply it by a team of fifty and it stops being a fun project and becomes somebody's full-time job, done badly, forever.
That is the honest case for buying rather than building, and it is a large part of why I am comfortable marketing Upsun Dispatch. I have now personally assembled the worse version. It was a wonderful way to spend nineteen evenings and it is not something I would hand to fifty engineers and ask them to depend on.
ask me about it
I am not selling this, and as the section above says, I would talk most people out of building one. It exists because I wanted to understand agentic delivery from the inside, on my own time, and now I do.
If you want to talk about what actually breaks in this stuff, or about the product work it feeds, I am easy to find.