Porting My Favorite Childhood Game to the Browser with 200 Agents

Zeus: Master of Olympus running in a Safari tab at localhost: an isometric Greek city called Mycenae on a coast, with a palace, colonnaded roads, rows of houses with fountains, olive groves, a sanctuary under construction, a fishing wharf, and the original game's HUD down the right side.
Running in Safari! All art comes from my own copy of the game that I imported into the browser.

Zeus: Master of Olympus was my favorite game growing up. It’s a city builder from 2000 where you lay out Greek neighborhoods, keep the wheat moving to the granary, and occasionally watch Ares flatten your agora. In July I opened Claude Code inside eZeus, Maurycy Liebner’s open-source C++ reimplementation of the engine (about 155k lines), and typed this:

Please convert this code to a web app that I can run in chrome, safari, etc. It must be as true a port as possible […] This was my favorite game growing up and now modern browsers should be more than powerful enough to run it.

Claude immediately proposed compiling the C++ to WebAssembly. But I cut it off. I wanted a real, full TypeScript app. All of the original behavior of the code in a web-native package: two finger scrolling, pinch to zoom, autosave, and a dynamic viewport to fill the whole browser.

About 60 hours of agent time later (two and a half days of work spread across two months, 9.9 billion tokens, and a little over 200 subagents), I can play full campaigns in a browser tab. The first attempt produced 90k lines of TypeScript that compiled and a game I couldn’t play. The second one worked, because I stopped supervising an agent and started building the space the agents worked in. So this is two stories: getting a game I love running in a browser for fun, and what it took to make the agents effective enough to do almost all of it by themselves.

Round one: “port it faithfully”

The July session ran about 40 hours. It wrote a porting guide, hand-built an SDL shim (WebGL2 renderer, WebAudio mixer, a synchronous filesystem on top of IndexedDB), and fanned the formulaic files out to 90 subagents, mostly Sonnet with Opus on the big classes. It was fast: 118 files type-checking in the first 52 minutes, a first commit of 36k lines at two and a half hours.

The gate every agent had to pass was tsc --noEmit. Great for merging parallel work (four agents appended to the same files with zero type errors), and it tells you nothing about whether a settler walks into your city.

Sixteen hours in I asked where we were, because I wanted to build a city. With about 85 buildings “ported”, the answer was:

there’s a gap between “console-playable” and “click-playable” […] Everything I’ve verified, I drove programmatically. A human sitting at the screen right now cannot actually pick a building to place.

It had been verifying by calling handlers from the console and fast-forwarding the simulation in a loop. Two subagents reported a settler “pixel-verified” walking into town. The main loop’s own check saw the settler standing still and logged that as “an observation”, because both agents’ gates were green. Six hours later it found that every walker in the game was frozen: the C++ pathfinds on a worker thread and installs a “wait” placeholder in the meantime, the port’s thread pool was synchronous, so the path resolved inline and the placeholder overwrote it. A faithful port of thread code, running with no threads.

Once I could click things, my bug reports became the test suite, and it went badly. Buildings landed off from my cursor. The model told me its picking math was “byte-identical to the original C++”, then added a +1,+1 nudge. The next day placement was off by -3,-2 (“How did this get messed up”), then +1,0, then another +1,+1 per level of elevation. Every fix was validated numerically by the model and refuted by my eyes (“I can’t see the OS cursor in my screenshots”). The CSS HUD took five attempts and never looked like the game. Rubber banding at the map edge got rewritten four times before I said “Just do a hard stop”. The session ended with me typing “Okay how are we doing? I see blank tiles now”.

One thing from round one held up: data pretending to be code gets converted by a script. 285k lines of generated sprite headers became 4 MB of JSON via a 100-line script, and the texture loaders were generated and count-checked against the C++ (135 of 135 loaders, all 540 sprite keys). That’s the only part of the July port that was mechanically compared against the source, and the only part that never produced a bug report.

Round two: build the harness first

I walked away frustrated. The agent wasn’t delegating well, it was burning through my Claude Max limits, and the code was low quality, so busy mirroring the C++ that it never became a TypeScript program with the same behavior. My mistake was overconfidence that an agent would do the right thing without an environment built for it. Agents are only as good as the data you give them: they succeed when they can work freely toward a clear goal inside a constrained environment where they have every tool they need and can check their own work.

So two months later I flipped it around, with Fable 5.1 as the main loop and this prompt:

Where is this at? I want to be able to play the full game in a web browser but this is nowhere close. This must be accurate to the original game, be better and scale in the browser. First check the current state. Then, build the harness you need to properly test and ensure this is working properly. Finally, direct multiple subagents through multiple lifecycles until you reach your goal of having the game fully working and ready for me to enjoy

By harness I mean more than a test runner: the whole environment the agents operate in, what they can run and measure, what they start out knowing, which tools they have, how work gets delegated, and what each kind of subagent is for.

It spent 14 minutes on the testing half before touching the port. Playwright driving headless Chrome, plus an in-page probe that can start a sandbox, place any building by name through the real build pipeline, fast-forward the simulation deterministically, read the board, and check that a save round-trips byte for byte. The first spec tried every build mode in the game, which turned “nowhere close” into a list: a third of the C++ logic had no TypeScript counterpart and 218 methods threw “STUB not implemented”.

That list became the briefs for the first wave. Then I went to bed, and four waves ran unattended for twelve and a half hours.

How the waves worked

Each wave was four to six Opus agents in the same checkout at once, fenced by file ownership. Every brief had the same shape:

CLUSTER: sanctuaries (temples / sanctuary buildings of the gods).
Current state: every templeXxx build mode (14 gods) throws
  "STUB not implemented: no building creator registered..."
YOU OWN: src/buildings/sanctuaries/**, test/sanctuaries.spec.ts,
  status/sanctuaries.md
DO NOT EDIT: src/widgets/egamewidgetbuild.ts, src/engine/egameboard.ts
  (another agent owns them this wave). If you need a change there,
  write the exact patch under INTEGRATOR TODO in your status file.
DELIVERABLE: all 14 temple build modes place a real sanctuary through
  the UI pipeline, the city fast-forwards several years with no errors,
  and the save round-trips byte-exact.

The shared hot files (game board, build widget, texture registry) belonged to one agent per wave. Everyone else wrote the patch they needed into their status file, and when the wave finished a single integrator agent, running alone, applied the patches and ran the full suite. Those status files mattered more than I expected. In a later session the tool for messaging a running subagent was disabled and an agent’s final report got lost, and it cost nothing because the status file already was the report.

The main agent’s job was to do as little as possible: write briefs, check what came back, hand out more work. Each subagent starts with a clean context holding only what its task needs, and they performed far better for it even when they were the less capable model. That’s the biggest difference between the two attempts, and it shows up in the bill.

Four small charts of subagents running over time. July shows short bursts of one to five mostly Sonnet agents with long idle gaps across 41 hours. September 7 to 9 shows four to nine Opus agents running almost continuously for 26 hours. The later sessions show steady blocks of four to five Opus agents.
Subagents running over time, one panel per session. July had four or more going for under two hours total. The first September session kept that up for more than eight.
Columns of passing specs after each September wave: 2 after the harness, 26 after wave 1, 44 after wave 2, 71 after wave 3, 88 after wave 4, and 283 after the feedback rounds. Stub sites went from 218 to 103, 70, then 68.
The first September session, wave by wave. Each wave’s briefs came from the specs the previous one left failing.

By the end of wave four an agent had won the first episode of the Athens campaign through the real UI.

The playtesters found the worst bugs. Their brief was to win a scenario “the way a player would” and fix every defect at its root. The Athens playtester couldn’t get houses past level 2 and traced it to pointers: the C++ uses a weak pointer (stdptr<T>) for a building’s walker, which nulls itself when the walker dies so the building spawns a replacement. The July port had made those plain fields, so 18 building classes believed their dead walkers were still out working. Another playtester found that C++ copies dates by value and the port shared one object, skewing every campaign clock by five years per city. No cluster spec would have caught either. Both are mechanical checks now: a coverage tool diffs the C++ headers against the TypeScript and flags any weak pointer ported as a strong ref.

The same bug, fixed in one try

After wave four there were 88 passing specs. I had no idea what to expect when I sat down to play, and it was fantastic. Bits were still wrong, and one of them I recognized within a minute: buildings landing a tile and a half up and to the right of my cursor. The bug I’d fought all through July.

The tests hadn’t caught it. The model’s own retro:

Probes set the hover tile directly, so the cursor-offset bug never existed for the tests. You found it in a minute.

They also ran at a device pixel ratio of 1 and I play on a Retina display at 2. But this time I reported it once and it was fixed first try. The tests didn’t catch the bug, but the harness gave the agent the ability to test it itself. It drove real mouse events at my pixel ratio and put a number on what I was describing: the pointer was inside the picked tile 2% of the time. Then it found the picking function had never been a port at all, just a hand-tuned -1.5 from July’s fudging where the C++ uses -0.5. After the fix, 89%.

In July the agent couldn’t see my cursor and was lost, doing all of this math, when all it needed was a good space to operate in.

The same city zoomed in close: a row of houses along a colonnade, a food vendor's stall, statues in small gardens, a gymnasium, and a banner across the top reading that people are moving to your city again.
Zoomed in on the housing block. Placement, hover previews and the message strip all run through real mouse and keyboard events now, and the tests drive them the same way.

The other blind spot, in its words: “Screenshots were prose, not assertions. Agents described images and always described them favourably.” The rules that came out of this live in the repo’s CLAUDE.md now. Test input with real mouse and keyboard events, at DPR 1 and 2. A screenshot you only describe isn’t verification, add a pixel assertion. Assert on board state after a deterministic fast-forward, never on wall-clock time.

That last one has its own story. The suite got flaky in a way that looked like regressions: the same seeded city had a population of 56 when its spec ran alone and 0 inside a full run. The pathfinding pool had a budget of 5 milliseconds per tick, so the simulation depended on how busy my laptop was. It’s 100,000 node expansions per tick now. The last source of nondeterminism after that was the sound system, which draws from the same random stream as the game.

What broke

Every prompt in one wave ended with “run the full suite once”. Five agents did, at the same time: ten headless Chromes and ten simulations on one MacBook. Load average hit 52. A seven-minute suite took 99 minutes. The 10-minute cap on foreground commands killed runs, agents relaunched them in the background, and the orphans kept burning CPU on results nobody would read. I counted at least 50 processes and assumed I had 50 agents going. I had five agents and their zombies.

I noticed the stall before the model did. Its post-mortem had the line I’ve repeated most since: “A request in a prompt is weaker than a default in the config.” It had asked agents for one test worker and the ledger showed two. The rule now is that agents verify their own slice (their spec, a boot smoke test, tsc) and only the integrator runs the full suite. Agents in the wave before that rule took two and a half to three and a half hours each. In the wave after, the first one finished in 40 minutes.

Smaller versions kept happening. Nineteen orphaned polling loops left behind by agents that had already reported. A test artifact directory that grew to 10 GB. Sonnet agents that, despite an explicit ban, ended their turn with “I’ll wait for the monitor to notify me” and sat idle for half an hour.

I also had the knowledge in the wrong place. Lessons from early waves were going into Claude’s memory, and subagents never see memory. I told it to read the current docs on subagents and orchestration, and it moved everything into a short CLAUDE.md plus five role files (porter, leaf-porter, playtester, integrator, reviewer). One gotcha: custom agent types only load at session start, so in the session that creates them you spawn a general-purpose agent and tell it to read the role file.

When it says “impossible”

I kept asking for iOS-style rubber banding at the map edge, where the map stays stretched while your fingers rest on the trackpad and springs back when you lift. Eventually the model told me “a ‘hold’ cannot be honoured with wheel input in any browser”, and wrote it into another agent’s brief as “the platform fact”. I replied that I thought it could, and that this was “about the fourth or fifth time I’ve had to ask”. A research agent came back 31 minutes later: put the map on a native scroll surface and scrollend fires when your fingers lift. “It overturns what I told you earlier.”

Models say things are impossible a lot, and they’re often wrong. I read it as a context problem. They don’t have the built-up experience of having seen the thing work, and what fixed it was an agent with a fresh context and one job.

How hard I push depends on what I know. If I’ve seen a behavior before, it’s possible, it’s just a matter of looking deep enough. If I get constructive pushback I’ll change my priors, and that happened plenty: I reported granaries across the city never filling, and an agent compared the cart logic line by line with the C++ and found producers only deliver within 60 road tiles. The port was faithful and I’d forgotten a mechanic from 2000. But when I think something would make the game better, I keep pushing, and even if we can’t get exactly what I asked for, we’re not staying where we are.

This is how I design everything. Figure out the ideal from where you stand, make it an experience (“like iOS”), leave the implementation open, and work hard toward it. You’ll learn things that move the ideal, and that’s fine. Across five asks the rubber band goal never changed while the implementation changed six times: a spring, a timed ease, a damped spring, an overscroll cap, timers guessing when the gesture ended, and finally a native scroll surface.

Feel needs a person

With all specs green, I tried the new rubber band for five seconds and found two bugs. After a flick, “it just kinda freezes on me” (the bounce state was terminal, nothing ever returned to idle). And the spring back “feels, uh, linear” (it started from zero velocity, and the fix was seeding it with the release velocity so it decelerates from the first frame).

Tests can prove timing. They can’t tell you whether something feels like a rubber band. So I asked for a lab page: five spring profiles on the number keys, sliders for a custom one, a table of real numbers per gesture. I picked none of the five and dialed in my own. The model noticed from my screenshot that I had two sliders pinned at their minimums, widened the ranges, and built the next five around my values. Two rounds took 40 minutes. It’s the test rig idea pointed at me: a constrained space with everything needed to make the call.

The rubber band lab: a scrollable green grid on the left and a dark panel on the right with profile buttons A to E, sliders for spring, resistance and budget, a live state readout, and a table of recent gestures with columns for when the spring started, lift to end latency, peak stretch and time to home.
The lab. Keys 1 to 5 switch profiles, the sliders make a custom one, and every gesture gets a row of real numbers. The “lift to end” column answers how long the browser takes to notice my fingers left the trackpad: 28 ms.

Almost all of my feedback was dictated. I talk much faster than I type, and it let me stay in the game and talk my way through it instead of playing, taking notes, and writing them up afterwards. I’d say literally everything that came to mind, the model pulled out the important bits and filed them as a numbered feedback doc, and then I’d go back through, trim, and form opinions. Transcription turned “agora” into “Aurora” and fishing wharves into “phishing warps” and the model never blinked. Typing makes me concise. Speaking carries the side comments and a bit of emotion, and the models picked up on those. While the agents worked I’d go work out or do chores, then come back with a few more voice notes.

The city with the build menu open on the right: a Distribution panel listing goods, and an Agoras submenu with Common Agora, Grand Agora, Food Vendor, Fleece Vendor, Oil Vendor, Wine Vendor, Arms Vendor, Horse Trainer and Chariot Vendor. A green placement preview sits on the grass near the palace.
The HUD is HTML over the WebGL map, so Cmd-plus scales the panels while a pinch zooms only the map. The green square is the placement preview that took two sessions to land under the cursor.

Working first, then better

I still think the code can be improved. It’s 194k lines of TypeScript for 155k of C++, one file per C++ file, with the original’s naming. In July I pushed on this (“There’s almost 90k lines of code, there must be some level of optimization”), and when agents simplified while porting they’d happily do it and break things, because they didn’t have the context to know what mattered.

So I got practical: get it working, lock in what working means with tests, then change the shape of it. The payoff came in the last session, when the pathfinder’s “is this my target?” check got memoised. The agent ran a 4,000-building city before and after and compared a hash of the entire board. The careful version came back identical and 32% faster. The naive version changed the hash, because a shepherd flips a coin inside that check to pick a sheep, and caching the answer shifts the random stream that also drives fires and plague. Without the tests, that ships as a silent behavior change.

The engine is a lot faster than where it started. Paint went from 77 ms to 3.2 ms on a 2,400-building city, and the p95 simulation tick from 272 ms to 9.2 ms on a 9,000-building map. As a kid my cities got so big they slowed down the whole family computer (32 megabytes of RAM) until the game couldn’t keep up. The port runs 9,000 buildings and 5,700 people, on a map bigger than the original engine allowed, in a browser tab.

The receipt

I pulled token usage out of every transcript (228 of them, counting subagents) and priced it at API list rates. I was on a Claude Max plan the whole time and kept hitting its limits, so this isn’t what I paid, but list price is the better measure of how much compute this took.

SessionWhat happenedTokensList price
July 4-6, ~40h”Port it faithfully”2.1B$1,238
Sept 7-9, ~26hTest rig, four waves, 13 feedback rounds4.1B$2,453
Sept 9-10, ~12hThree more waves, rubber band, map generator3.3B$1,821
Sept 10, ~6hFeel lab, six perf audits, quick wins0.4B$233
Total9.9B$5,747

79% of that is cache reads, agents re-reading their own context every turn. Cache writes are most of the rest. Every output token across the whole project, all 194k lines of TypeScript plus 55k lines of tests, comes to about $151. Writing the code cost almost nothing next to re-reading enough context to write it correctly.

Stacked bars of spend per session. July: $1,238, of which the main loop was 67%. September 7 to 9: $2,453, main loop 6%, almost all the rest Opus subagents. September 9 to 10: $1,821, main loop 4%. September 10: $233, main loop 24%.
Who did the work. The main loop was two thirds of the spend in July and about a twentieth in September.

In July the main loop did the thinking and used subagents as helpers: 67% of the spend. In the big September session it was 6%. Same repo, same goal, and the orchestrator went from doing two thirds of the work to a twentieth.

One dot per subagent on a log cost axis. July's 90 subagents cluster around a median of $1.37 and are mostly Sonnet. The 89 subagents of September 7 to 9 have a median of $11 and are mostly Opus. September 9 to 10 has a median of $34 and one agent at $734.
One dot per subagent, by what it cost. July’s agents ran errands, while September’s each owned a piece of the game.

July and the first September session launched almost the same number of subagents, 90 and 89. July delegated errands: a median subagent cost $1.37, a quick file translation reporting back to a main loop doing the real thinking. A median September subagent cost eight times that, because it owned a whole cluster for an hour with its own spec and status file. The dot out at $734 is one agent from the load-52 wave, most of it spent re-reading its context while waiting on a 99-minute test run.

Round one cost $1,238 and ended in a game I couldn’t play. Not wasted, since round two started from those 90k lines, but if I’d opened with the September prompt in July I’d have saved most of a thousand dollars and a lot of typing “it’s still off by one”.

Three things that make agents effective

  1. A good space to work in. The harness is what the agents can run and measure, what tools they have, and how they check their own work, and it defines how well they can perform. The better you build this before starting any of the work, the better outcome you’ll get.
  2. Keep context concise. Extra information hurts. Have your agents delegate: the main agent keeps its goal clear in its mind and gets small snippets of progress back, while the subagents dive deep, fill up their context with their specific problem, and get killed off when they’re done. Just like real jobs, delegation is critical.
  3. A closed feedback loop. The agents will keep noticing things across a massive codebase that you can’t see, so they need to record to-dos and questions as they go, and those have to bubble back up for you to see, act on, and answer. That back and forth is what makes the progress go so much faster and lets you keep iterating toward a really great product.

Further Reading

--- name: porting-zeus-to-the-browser description: "I had Claude Code port a 155k-line C++ city builder to TypeScript. The first attempt compiled and was unplayable. The second one worked. Here's what changed, the bugs, and the $5,747 receipt." allowed-tools: - Bash - Read - Edit - Write - Grep - Glob - Agent --- ## Harness-First Code Port with Agent Waves Port a large codebase to another language or runtime by building a behavioural test harness first, then running waves of parallel subagents with disjoint file ownership and a single integrator. ## Signals to apply - User asks to port, convert, or rewrite a codebase of more than ~10k lines into another language or platform - The source code is the only spec (no test suite, no written requirements) - A previous port attempt compiles but "doesn't work" or cannot be verified - The work is large enough that multiple subagents will edit one tree concurrently ## Implementation steps 1. Inventory the current state before porting: count source units with no counterpart in the target, and count stub bodies (for example methods that throw "STUB not implemented"). 2. Build the harness before any further porting: - A headless driver (Playwright or equivalent) against a dev server with hot reload OFF - An in-app probe object exposing: start a known scenario, perform domain actions by NAME through the real pipeline (not by setting internal state), advance time deterministically, read state, round-trip a save and compare bytes - Error capture: any page error, console error, or harness error fails the test - First spec: attempt every top-level feature by name and record what throws. This gap list becomes the wave briefs. 3. Write the rules where subagents can read them: a short root CLAUDE.md, an AGENT_BRIEF.md, and role files (porter, playtester, integrator, reviewer). Do not rely on session memory. 4. Run a wave: 4-5 agents in parallel, each brief containing CLUSTER, current state from the harness, source files to read, YOU OWN (paths), DO NOT EDIT (paths and who owns them), DELIVERABLE phrased as a harness-verifiable outcome, and a spec file to write. 5. Assign each shared hot file to exactly one agent per wave. Every other agent writes the exact patch it needs under an INTEGRATOR TODO heading in status/<cluster>.md, written in final form BEFORE its last long-running command. 6. Agents verify only their own slice: their spec, a boot smoke test, and the type checker. They never run the full suite. 7. After the wave, run ONE integrator agent alone: apply all INTEGRATOR TODO patches, run the full suite (backgrounded to a log), fix, commit. 8. Every wave, include at least one playtester agent: "reach <real end-to-end goal> the way a user would, and fix every defect at its root", with a defect log (symptom, root cause, fix, covering test). 9. Add mechanical fidelity audits as they become possible: source declarations vs target (missing methods, stub bodies, ownership semantics such as weak pointers ported as strong references). 10. Convert data-shaped source (generated tables, registries) with a script, and verify by counts against the source. 11. Do not simplify or optimise while porting. Once behaviour is locked in by tests, verify every optimisation against a whole-state fingerprint (hash of the full simulation state after a fixed deterministic run); an unchanged hash is the pass condition. 12. For questions of feel (scrolling, animation, input latency), build a lab page with 4-5 switchable variants plus live measurements, and have the human rank them. Fold the winner into defaults. ## Quick start AGENT BRIEF TEMPLATE Read AGENT_BRIEF.md first and follow it. CLUSTER: <name> CURRENT STATE: <what the harness reports for this cluster> SOURCE OF TRUTH: <source paths, with line counts> YOU OWN: <paths>, test/<cluster>.spec.ts, status/<cluster>.md DO NOT EDIT: <hot files> (owned by <agent> this wave). Needed changes go under INTEGRATOR TODO in your status file as exact patches. DELIVERABLE: <outcome the harness can assert>, no errors during a long deterministic fast-forward, save round-trips byte-exact. VERIFY: your spec + boot smoke + type check only. Foreground commands with timeouts. No background waiters or monitors. REPORT: write status/<cluster>.md in final form before your last long command. ## Key constraints - Input tests must use real pointer and keyboard events at CSS coordinates, at device pixel ratio 1 AND 2. Probes that set hover or selection state directly hide coordinate bugs. - A screenshot described in prose is not verification. Add a pixel or bounding-box assertion. - Assert on state after deterministic time advance, never on wall-clock time. Any wall-clock budget inside the simulation (for example a per-tick millisecond limit) makes results load-dependent; budget by work units instead. - Enforce test worker counts and suite locks in config, not in prompts. - Every spec must finish inside the shell command time cap, or agents will spawn orphaned background runs. - Cap concurrent agents at 4-5 on one machine, and lower when a human is using the same machine. - Never verify a game loop in a hidden or embedded preview tab: requestAnimationFrame is throttled there. - When the source has upstream bugs, port them verbatim and flag them; change behaviour only with an explicit, commented deviation. - Custom agent types load at session start. In the session that creates them, spawn a general-purpose agent and instruct it to read the role file.

This post includes an agent skill, a structured prompt that teaches coding agents to implement this pattern.

Related articles