Porting My Favorite Childhood Game to the Browser with 200 Agents

Zeus: Master of Olympus was my favorite game growing up. It’s a city builder from 2000 where you lay out Greek neighborhoods, keep the wheat moving to the granary, and occasionally watch Ares flatten your agora. In July I opened Claude Code inside eZeus, Maurycy Liebner’s open-source C++ reimplementation of the engine (about 155k lines), and typed this:
Please convert this code to a web app that I can run in chrome, safari, etc. It must be as true a port as possible […] This was my favorite game growing up and now modern browsers should be more than powerful enough to run it.
Claude immediately proposed compiling the C++ to WebAssembly. But I cut it off. I wanted a real, full TypeScript app. All of the original behavior of the code in a web-native package: two finger scrolling, pinch to zoom, autosave, and a dynamic viewport to fill the whole browser.
About 60 hours of agent time later (two and a half days of work spread across two months, 9.9 billion tokens, and a little over 200 subagents), I can play full campaigns in a browser tab. The first attempt produced 90k lines of TypeScript that compiled and a game I couldn’t play. The second one worked, because I stopped supervising an agent and started building the space the agents worked in. So this is two stories: getting a game I love running in a browser for fun, and what it took to make the agents effective enough to do almost all of it by themselves.
Round one: “port it faithfully”
The July session ran about 40 hours. It wrote a porting guide, hand-built an SDL shim (WebGL2 renderer, WebAudio mixer, a synchronous filesystem on top of IndexedDB), and fanned the formulaic files out to 90 subagents, mostly Sonnet with Opus on the big classes. It was fast: 118 files type-checking in the first 52 minutes, a first commit of 36k lines at two and a half hours.
The gate every agent had to pass was tsc --noEmit. Great for merging parallel work (four agents appended to the same files with zero type errors), and it tells you nothing about whether a settler walks into your city.
Sixteen hours in I asked where we were, because I wanted to build a city. With about 85 buildings “ported”, the answer was:
there’s a gap between “console-playable” and “click-playable” […] Everything I’ve verified, I drove programmatically. A human sitting at the screen right now cannot actually pick a building to place.
It had been verifying by calling handlers from the console and fast-forwarding the simulation in a loop. Two subagents reported a settler “pixel-verified” walking into town. The main loop’s own check saw the settler standing still and logged that as “an observation”, because both agents’ gates were green. Six hours later it found that every walker in the game was frozen: the C++ pathfinds on a worker thread and installs a “wait” placeholder in the meantime, the port’s thread pool was synchronous, so the path resolved inline and the placeholder overwrote it. A faithful port of thread code, running with no threads.
Once I could click things, my bug reports became the test suite, and it went badly. Buildings landed off from my cursor. The model told me its picking math was “byte-identical to the original C++”, then added a +1,+1 nudge. The next day placement was off by -3,-2 (“How did this get messed up”), then +1,0, then another +1,+1 per level of elevation. Every fix was validated numerically by the model and refuted by my eyes (“I can’t see the OS cursor in my screenshots”). The CSS HUD took five attempts and never looked like the game. Rubber banding at the map edge got rewritten four times before I said “Just do a hard stop”. The session ended with me typing “Okay how are we doing? I see blank tiles now”.
One thing from round one held up: data pretending to be code gets converted by a script. 285k lines of generated sprite headers became 4 MB of JSON via a 100-line script, and the texture loaders were generated and count-checked against the C++ (135 of 135 loaders, all 540 sprite keys). That’s the only part of the July port that was mechanically compared against the source, and the only part that never produced a bug report.
Round two: build the harness first
I walked away frustrated. The agent wasn’t delegating well, it was burning through my Claude Max limits, and the code was low quality, so busy mirroring the C++ that it never became a TypeScript program with the same behavior. My mistake was overconfidence that an agent would do the right thing without an environment built for it. Agents are only as good as the data you give them: they succeed when they can work freely toward a clear goal inside a constrained environment where they have every tool they need and can check their own work.
So two months later I flipped it around, with Fable 5.1 as the main loop and this prompt:
Where is this at? I want to be able to play the full game in a web browser but this is nowhere close. This must be accurate to the original game, be better and scale in the browser. First check the current state. Then, build the harness you need to properly test and ensure this is working properly. Finally, direct multiple subagents through multiple lifecycles until you reach your goal of having the game fully working and ready for me to enjoy
By harness I mean more than a test runner: the whole environment the agents operate in, what they can run and measure, what they start out knowing, which tools they have, how work gets delegated, and what each kind of subagent is for.
It spent 14 minutes on the testing half before touching the port. Playwright driving headless Chrome, plus an in-page probe that can start a sandbox, place any building by name through the real build pipeline, fast-forward the simulation deterministically, read the board, and check that a save round-trips byte for byte. The first spec tried every build mode in the game, which turned “nowhere close” into a list: a third of the C++ logic had no TypeScript counterpart and 218 methods threw “STUB not implemented”.
That list became the briefs for the first wave. Then I went to bed, and four waves ran unattended for twelve and a half hours.
How the waves worked
Each wave was four to six Opus agents in the same checkout at once, fenced by file ownership. Every brief had the same shape:
CLUSTER: sanctuaries (temples / sanctuary buildings of the gods).
Current state: every templeXxx build mode (14 gods) throws
"STUB not implemented: no building creator registered..."
YOU OWN: src/buildings/sanctuaries/**, test/sanctuaries.spec.ts,
status/sanctuaries.md
DO NOT EDIT: src/widgets/egamewidgetbuild.ts, src/engine/egameboard.ts
(another agent owns them this wave). If you need a change there,
write the exact patch under INTEGRATOR TODO in your status file.
DELIVERABLE: all 14 temple build modes place a real sanctuary through
the UI pipeline, the city fast-forwards several years with no errors,
and the save round-trips byte-exact.The shared hot files (game board, build widget, texture registry) belonged to one agent per wave. Everyone else wrote the patch they needed into their status file, and when the wave finished a single integrator agent, running alone, applied the patches and ran the full suite. Those status files mattered more than I expected. In a later session the tool for messaging a running subagent was disabled and an agent’s final report got lost, and it cost nothing because the status file already was the report.
The main agent’s job was to do as little as possible: write briefs, check what came back, hand out more work. Each subagent starts with a clean context holding only what its task needs, and they performed far better for it even when they were the less capable model. That’s the biggest difference between the two attempts, and it shows up in the bill.
By the end of wave four an agent had won the first episode of the Athens campaign through the real UI.
The playtesters found the worst bugs. Their brief was to win a scenario “the way a player would” and fix every defect at its root. The Athens playtester couldn’t get houses past level 2 and traced it to pointers: the C++ uses a weak pointer (stdptr<T>) for a building’s walker, which nulls itself when the walker dies so the building spawns a replacement. The July port had made those plain fields, so 18 building classes believed their dead walkers were still out working. Another playtester found that C++ copies dates by value and the port shared one object, skewing every campaign clock by five years per city. No cluster spec would have caught either. Both are mechanical checks now: a coverage tool diffs the C++ headers against the TypeScript and flags any weak pointer ported as a strong ref.
The same bug, fixed in one try
After wave four there were 88 passing specs. I had no idea what to expect when I sat down to play, and it was fantastic. Bits were still wrong, and one of them I recognized within a minute: buildings landing a tile and a half up and to the right of my cursor. The bug I’d fought all through July.
The tests hadn’t caught it. The model’s own retro:
Probes set the hover tile directly, so the cursor-offset bug never existed for the tests. You found it in a minute.
They also ran at a device pixel ratio of 1 and I play on a Retina display at 2. But this time I reported it once and it was fixed first try. The tests didn’t catch the bug, but the harness gave the agent the ability to test it itself. It drove real mouse events at my pixel ratio and put a number on what I was describing: the pointer was inside the picked tile 2% of the time. Then it found the picking function had never been a port at all, just a hand-tuned -1.5 from July’s fudging where the C++ uses -0.5. After the fix, 89%.
In July the agent couldn’t see my cursor and was lost, doing all of this math, when all it needed was a good space to operate in.

The other blind spot, in its words: “Screenshots were prose, not assertions. Agents described images and always described them favourably.” The rules that came out of this live in the repo’s CLAUDE.md now. Test input with real mouse and keyboard events, at DPR 1 and 2. A screenshot you only describe isn’t verification, add a pixel assertion. Assert on board state after a deterministic fast-forward, never on wall-clock time.
That last one has its own story. The suite got flaky in a way that looked like regressions: the same seeded city had a population of 56 when its spec ran alone and 0 inside a full run. The pathfinding pool had a budget of 5 milliseconds per tick, so the simulation depended on how busy my laptop was. It’s 100,000 node expansions per tick now. The last source of nondeterminism after that was the sound system, which draws from the same random stream as the game.
What broke
Every prompt in one wave ended with “run the full suite once”. Five agents did, at the same time: ten headless Chromes and ten simulations on one MacBook. Load average hit 52. A seven-minute suite took 99 minutes. The 10-minute cap on foreground commands killed runs, agents relaunched them in the background, and the orphans kept burning CPU on results nobody would read. I counted at least 50 processes and assumed I had 50 agents going. I had five agents and their zombies.
I noticed the stall before the model did. Its post-mortem had the line I’ve repeated most since: “A request in a prompt is weaker than a default in the config.” It had asked agents for one test worker and the ledger showed two. The rule now is that agents verify their own slice (their spec, a boot smoke test, tsc) and only the integrator runs the full suite. Agents in the wave before that rule took two and a half to three and a half hours each. In the wave after, the first one finished in 40 minutes.
Smaller versions kept happening. Nineteen orphaned polling loops left behind by agents that had already reported. A test artifact directory that grew to 10 GB. Sonnet agents that, despite an explicit ban, ended their turn with “I’ll wait for the monitor to notify me” and sat idle for half an hour.
I also had the knowledge in the wrong place. Lessons from early waves were going into Claude’s memory, and subagents never see memory. I told it to read the current docs on subagents and orchestration, and it moved everything into a short CLAUDE.md plus five role files (porter, leaf-porter, playtester, integrator, reviewer). One gotcha: custom agent types only load at session start, so in the session that creates them you spawn a general-purpose agent and tell it to read the role file.
When it says “impossible”
I kept asking for iOS-style rubber banding at the map edge, where the map stays stretched while your fingers rest on the trackpad and springs back when you lift. Eventually the model told me “a ‘hold’ cannot be honoured with wheel input in any browser”, and wrote it into another agent’s brief as “the platform fact”. I replied that I thought it could, and that this was “about the fourth or fifth time I’ve had to ask”. A research agent came back 31 minutes later: put the map on a native scroll surface and scrollend fires when your fingers lift. “It overturns what I told you earlier.”
Models say things are impossible a lot, and they’re often wrong. I read it as a context problem. They don’t have the built-up experience of having seen the thing work, and what fixed it was an agent with a fresh context and one job.
How hard I push depends on what I know. If I’ve seen a behavior before, it’s possible, it’s just a matter of looking deep enough. If I get constructive pushback I’ll change my priors, and that happened plenty: I reported granaries across the city never filling, and an agent compared the cart logic line by line with the C++ and found producers only deliver within 60 road tiles. The port was faithful and I’d forgotten a mechanic from 2000. But when I think something would make the game better, I keep pushing, and even if we can’t get exactly what I asked for, we’re not staying where we are.
This is how I design everything. Figure out the ideal from where you stand, make it an experience (“like iOS”), leave the implementation open, and work hard toward it. You’ll learn things that move the ideal, and that’s fine. Across five asks the rubber band goal never changed while the implementation changed six times: a spring, a timed ease, a damped spring, an overscroll cap, timers guessing when the gesture ended, and finally a native scroll surface.
Feel needs a person
With all specs green, I tried the new rubber band for five seconds and found two bugs. After a flick, “it just kinda freezes on me” (the bounce state was terminal, nothing ever returned to idle). And the spring back “feels, uh, linear” (it started from zero velocity, and the fix was seeding it with the release velocity so it decelerates from the first frame).
Tests can prove timing. They can’t tell you whether something feels like a rubber band. So I asked for a lab page: five spring profiles on the number keys, sliders for a custom one, a table of real numbers per gesture. I picked none of the five and dialed in my own. The model noticed from my screenshot that I had two sliders pinned at their minimums, widened the ranges, and built the next five around my values. Two rounds took 40 minutes. It’s the test rig idea pointed at me: a constrained space with everything needed to make the call.

Almost all of my feedback was dictated. I talk much faster than I type, and it let me stay in the game and talk my way through it instead of playing, taking notes, and writing them up afterwards. I’d say literally everything that came to mind, the model pulled out the important bits and filed them as a numbered feedback doc, and then I’d go back through, trim, and form opinions. Transcription turned “agora” into “Aurora” and fishing wharves into “phishing warps” and the model never blinked. Typing makes me concise. Speaking carries the side comments and a bit of emotion, and the models picked up on those. While the agents worked I’d go work out or do chores, then come back with a few more voice notes.

Working first, then better
I still think the code can be improved. It’s 194k lines of TypeScript for 155k of C++, one file per C++ file, with the original’s naming. In July I pushed on this (“There’s almost 90k lines of code, there must be some level of optimization”), and when agents simplified while porting they’d happily do it and break things, because they didn’t have the context to know what mattered.
So I got practical: get it working, lock in what working means with tests, then change the shape of it. The payoff came in the last session, when the pathfinder’s “is this my target?” check got memoised. The agent ran a 4,000-building city before and after and compared a hash of the entire board. The careful version came back identical and 32% faster. The naive version changed the hash, because a shepherd flips a coin inside that check to pick a sheep, and caching the answer shifts the random stream that also drives fires and plague. Without the tests, that ships as a silent behavior change.
The engine is a lot faster than where it started. Paint went from 77 ms to 3.2 ms on a 2,400-building city, and the p95 simulation tick from 272 ms to 9.2 ms on a 9,000-building map. As a kid my cities got so big they slowed down the whole family computer (32 megabytes of RAM) until the game couldn’t keep up. The port runs 9,000 buildings and 5,700 people, on a map bigger than the original engine allowed, in a browser tab.
The receipt
I pulled token usage out of every transcript (228 of them, counting subagents) and priced it at API list rates. I was on a Claude Max plan the whole time and kept hitting its limits, so this isn’t what I paid, but list price is the better measure of how much compute this took.
| Session | What happened | Tokens | List price |
|---|---|---|---|
| July 4-6, ~40h | ”Port it faithfully” | 2.1B | $1,238 |
| Sept 7-9, ~26h | Test rig, four waves, 13 feedback rounds | 4.1B | $2,453 |
| Sept 9-10, ~12h | Three more waves, rubber band, map generator | 3.3B | $1,821 |
| Sept 10, ~6h | Feel lab, six perf audits, quick wins | 0.4B | $233 |
| Total | 9.9B | $5,747 |
79% of that is cache reads, agents re-reading their own context every turn. Cache writes are most of the rest. Every output token across the whole project, all 194k lines of TypeScript plus 55k lines of tests, comes to about $151. Writing the code cost almost nothing next to re-reading enough context to write it correctly.
In July the main loop did the thinking and used subagents as helpers: 67% of the spend. In the big September session it was 6%. Same repo, same goal, and the orchestrator went from doing two thirds of the work to a twentieth.
July and the first September session launched almost the same number of subagents, 90 and 89. July delegated errands: a median subagent cost $1.37, a quick file translation reporting back to a main loop doing the real thinking. A median September subagent cost eight times that, because it owned a whole cluster for an hour with its own spec and status file. The dot out at $734 is one agent from the load-52 wave, most of it spent re-reading its context while waiting on a 99-minute test run.
Round one cost $1,238 and ended in a game I couldn’t play. Not wasted, since round two started from those 90k lines, but if I’d opened with the September prompt in July I’d have saved most of a thousand dollars and a lot of typing “it’s still off by one”.
Three things that make agents effective
- A good space to work in. The harness is what the agents can run and measure, what tools they have, and how they check their own work, and it defines how well they can perform. The better you build this before starting any of the work, the better outcome you’ll get.
- Keep context concise. Extra information hurts. Have your agents delegate: the main agent keeps its goal clear in its mind and gets small snippets of progress back, while the subagents dive deep, fill up their context with their specific problem, and get killed off when they’re done. Just like real jobs, delegation is critical.
- A closed feedback loop. The agents will keep noticing things across a massive codebase that you can’t see, so they need to record to-dos and questions as they go, and those have to bubble back up for you to see, act on, and answer. That back and forth is what makes the progress go so much faster and lets you keep iterating toward a really great product.
Further Reading
This post includes an agent skill, a structured prompt that teaches coding agents to implement this pattern.