SteonMod
View as Markdown

Making fusions cheap for developers' AI agents

Written 2026-10-03 from research done that day (sources at the end). The user's goal: making a fusion in Fusion should cost a developer's AI agent far fewer tokens than today's hand-built mods, and Fusion should be the best at this. VISION.md §9 has the summary; BUILD.md steps 5.2, 5.5, 5.7 and 5.10 build it.

Where an agent's tokens go

An agent pays for three things:

  1. What it writes.
  2. What it reads: docs, source, tool output and logs. Everything it has read is re-sent on every turn, so it is paid again each turn (cheaply if cached).
  3. Turns spent stuck: each failed try adds more writing and more reading.

Today's passthrough mods are expensive on all three:

  • SkyCraft is about 890 KB of C++ and Java, and ValCraft's Valheim side alone is about 358 KB of C#.
  • Agents clone SkyCraft and read it to learn the design.
  • They grep huge dumps (UE4SS's SDK dump, decompiled C#) to find the player.
  • The guides name the usual places they get stuck: values landing in the wrong coordinate space, and the two games' frame clocks drifting apart.

What the research says

FindingSourceWhat it means for Fusion
Every token in a tool's answer is a token not used for reasoning. Concise answers by default, a detailed mode on request, paging and filtering, errors that say how to fix the call. One combined tool beats several small ones whose outputs must be chained. A concise answer took 72 tokens where the detailed one took 206.Anthropic, Writing effective tools for agentsEvery fusion command answers in a few lines by default, has --detailed, pages long lists, and gives fix hints in errors.
Improve tools by running an agent on 25+ realistic tasks and measuring tokens, tool calls, errors and success, then refining.sameThe router bench (method 10).
Agents should load context "just in time" through tools, not all up front. Long contexts are compacted.Anthropic, Effective context engineering for AI agentsSearch commands instead of reading dumps.
Model accuracy drops as input grows, long before the window is full, and related-but-irrelevant text hurts most.Chroma, Context Rot (18 models)Short pack, short outputs, short sessions.
Repository context files (AGENTS.md) did not improve success on average and raised cost by 20%+, because agents follow them and explore more. Human-written ones helped a little (~4%); LLM-written ones hurt. They're useful for non-standard rules.ETH Zurich, Evaluating AGENTS.md (2026)The pack holds only what an agent can't guess.
A second controlled study (Claude Code and Codex, 288 runs) found context files didn't measurably change correctness. Agents failed on design and exact wiring, not missing knowledge.Do Context Files Help Coding Agents? (2026)Fusion should remove the wiring (kit defaults), not explain it.
Guidance that is tuned by testing it on probe tasks raised the solve rate from 28.3% to 33.0%, mostly by helping agents find the right place faster.Probe-and-Refine Tuning of Repository Guidance (2026)Tune the pack with the bench; don't write it once and hope.
Skills load in steps: name and description first, the body when relevant, reference files when needed. One level of this helps when the material is large; a second level never helped and sometimes broke accuracy.Agent Skills standard (Dec 2025; read by Claude Code, Codex, Cursor, Gemini CLI…); Is Progressive Disclosure All You Need? (2026)One SKILL.md + flat reference files. No deeper tree.
Agents calling tools through code instead of one call at a time cut one workflow from 150,000 to 2,000 tokens. MCP servers load their tool descriptions up front; third-party benchmarks report CLIs using several to 30+ times fewer tokens for the same jobs.Anthropic, Code execution with MCP; MCP-vs-CLI benchmarks (2026, vendor-run, so treat numbers as rough)CLI first. An MCP server, if ever, is a thin wrapper over the same CLI.
94% of LLM compile errors were type errors; type checking cut them by 75%.ETH/Berkeley, Type-Constrained Code Generation (2025)Typed Router and Lua APIs, checked before anything runs.
The highest-leverage thing is a check the agent can run that returns pass or fail.Claude Code best practicesfusion router check, one command.
Cached prompt text costs 10% of normal; Claude Code, Codex and Cursor cache automatically, but only an unchanged prefix.Anthropic prompt cachingKeep the pack's text stable.
Multi-agent setups use about 15× the tokens of a chat (single agents about 4×).Anthropic, How we built our multi-agent research systemRecommend one agent per router.
universal-modder (the most-starred agent modding toolkit) uses a scan CLI, 12 engine playbooks as skills, and a shared knowledge base of per-game field notes so the next agent skips rediscovery.universal-modder READMEFusion keeps per-game notes with each router, searchable.

The method (ranked by how much it saves)

  1. Write almost nothing: kit defaults and data-first routers.
    • The kit fills every port it can from engine conventions:
      • Unreal: the player is player controller 0's pawn, the camera is its camera manager, collision comes from traces, characters are Character actors.
      • Unity: Camera.main, the player's CharacterController or Rigidbody, colliders.
    • A router states only what's different. Most of it is data in router.toml ("NPCs are class BP_Zombie_C", "health is field Health"), checked against a schema. Lua is for real exceptions, like getting past a main menu.
    • fusion router new --auto drafts the router from the running game. Router Studio's point-and-click writes the same data, at zero tokens.
    • Measured on what exists: about a third of VEIN's router (trace helpers, finding open ground, widget lookup) is generic Unreal code that belongs in the kit.
  2. Find, don't read.
    • fusion game find <text> and fusion game inspect <object> query the running game through the kit. They return the few best matches with readable names and short descriptions, page long results, and say how to narrow a cut-off list. No one greps a dump of tens of megabytes again.
    • Per-game notes live with each router (which class is what, version quirks, what didn't work), and fusion kb search <game> searches them. The second agent to touch a game pays nothing for discovery.
  3. Catch mistakes before running: types and schemas.
    • The Router API and the addon Lua API are published as type definition files. router check type-checks first: luau-lsp analyze --definitions for Luau, or LuaLS annotations and --check if Unreal router scripts stay on UE4SS's Lua 5.4. That makes typing a factor in step 5.2's choice; Luau has the stronger checker.
    • router.toml, scenarios and packages have a schema, so a wrong field comes back with its line number before anything starts.
  4. One command, short answers. fusion router check runs everything in order:
    1. types;
    2. manifest;
    3. hot-reload the router into the running game;
    4. conformance tests;
    5. picture checks reported as numbers ("layer: 97% non-empty, depth 0.4–180 m").

The default answer is a few lines, each failure with a hint (port pose: no data for 2 s. Router.player returned nil; try: fusion game find Pawn). --json is for scripts, and the exit code means pass or fail. Router Lua reloads without restarting the game, so a fix is tested in seconds.

  1. Test without the game: record and replay. fusion session record saves a real session's protocol traffic: frames, collision, entities, events. fusion replay plays it back into Fusion with no game running, the same every time. Experiences, addons, Devices and Fusion-side logic are tested this way in seconds, by an agent, with no human at the PC. Only router code needs the live game.
  2. Read little, and read it cheaply: the agent pack.
    • A skill folder per kit, one level deep. SKILL.md (under ~150 lines) holds the rules an agent can't guess (no game files, never get around anti-cheat, Lua only through ports), the loop (recon → draft → check → playtest) and the commands. Flat reference files hold the API and one example router.
    • The same text goes in AGENTS.md for tools without skills.
    • Only non-obvious content: everything else is either in the types or found by the commands.
    • The text stays stable (no changing dates at the top), so prompt caching serves it at a tenth of the price.
    • It's generated from the same source as the API docs, so it never goes stale.
  3. CLI, not MCP. A CLI costs nothing until it's used; MCP tool descriptions are paid on every session. If an MCP server is ever wanted, it wraps the same CLI.
  4. One agent, short sessions. The pack recommends one agent per router, a plan in MODLOG.md, and a fresh session with a short status note when the context grows. router new creates both files.
  5. Game patches: diff, don't rediscover.
    • fusion game diff compares the game's classes and fields between two builds and lists what was renamed or moved.
    • After a patch, router check names the port that broke.
    • Version lock is today's biggest ongoing cost; this turns it into a short fix.
  6. Measure and tune: the router bench.
    • A fixed set of tasks:
      • make a router for fake-guest variants;
      • rebuild VEIN's router from scratch;
      • fix a router broken by a simulated game update;
      • write an Experience;
      • add a block to an existing router.
    • Claude Code and Codex run them headless. The bench records tokens, turns, time and success.
    • Every change to the API, CLI output or pack is checked against the bench, so nothing gets worse, and tuned with it (probe-and-refine).
    • The first run sets the baseline. "Making a router costs N tokens" then becomes a number Fusion can publish and keep lowering.

The router bench: the baseline (2026-10-04)

Method 10 exists: tools/router-bench (its README says how it works). Five tasks run headless in fresh copies of Fusion and are graded by running the result. The Unreal tasks run in simulated Unreal games that run Fusion's real kit. The baseline is today's Fusion: the Router API (5.2), its docs, the example routers, fusion-cli, and nothing from methods 2-9 yet. One run per task and agent:

TaskClaude Code (Opus 5.5)Codex (gpt-6.1-sol, xhigh)
experience: an Experience with two goals, a timer, a lock137k tokens (12k fresh), 5 tool calls, 0.4 min156k (22k), 9, 1.1 min
add-block: a wildlife block in VEIN's router485k (36k), 14, 1.1 min196k (31k), 15, 1.2 min
unreal-update: VEIN's router after an update464k (36k), 12, 0.9 min528k (51k), 19, 2.3 min
unreal-new: Pinewood, a new Unreal game, from nothing554k (45k), 12, 1.2 min503k (65k), 17, 2.6 min
fake-variant: Arena, a fake-guest started with options868k (48k), 19, 2.2 min545k (66k), 19, 3.7 min

All 10 pass. Tokens are all input (cached too) plus output; "fresh" is input not read from the prompt cache. Claude Code's runs cost $0.17-0.69 at API prices.

What it shows:

  • The Router API already does most of the work. Pinewood's router from nothing is about 70 lines (data plus a few lines of Lua for its menu), made in one or two minutes.
  • Tokens are mostly context re-sent every turn. Claude Code starts at about 24k tokens before reading anything, and 90-95% of all input is cache reads. Each turn costs roughly the whole context again, so fewer turns (method 4) and less read early (method 6) are the levers.
  • What agents read: every run read docs/router-api.md and an example router (about 20k characters) and grepped the dump (10-15k). game find (method 2) replaces the greps; the pack replaces the reading.
  • Undocumented behavior is the most expensive thing. The smallest router, Arena's, cost the most for both agents: nothing says how a launch path for a game outside Steam is resolved, so both read Fusion's source (session.rs, fusion_paths.gd, fusion-cli). Claude ended with {tools}/../../games/arena/Arena.exe. Fusion needs a placeholder for a game's own folder, and the docs need one sentence.

Then the tools and the pack (2026-10-04, Claude, one run per task, tokens / tool calls):

TaskBaseline+ CLI tools (--label tools)+ pack (--label pack)
fake-variant (Arena)868k / 19403k / 10300k / 9
unreal-new (Pinewood)554k / 121,120k / 20774k / 18
unreal-update464k / 12308k / 10488k / 12
add-block485k / 14failed (a brittle test of ours)576k / 16
experience137k / 5137k / 5190k / 6

The tools: {game} in launch recipes, fusion-cli game, router check, router new. The pack: then kits/unreal/agent/SKILL.md (Unreal only) as a Claude Code skill and AGENTS.md; since 2026-10-04 (session 20) the pack is Fusion's agent/ folder for every path (SKILL.md plus unreal.md, unity.md, sdk.md and example routers of made-up games), measured with the same --pack. Every agent with the pack loaded it first and used router new and router check.

What it shows so far:

  • One run per task is too noisy. Differences of ±50% appear between runs of the same setup; only Arena's drop is clearly real. Next: --repeat 3 per setup.
  • Whatever agents read, keep small. Without the pack, the Pinewood agent read the kit's router_api.lua, which the new question code had made 40% bigger; it's split out now.
  • A tool used at the wrong moment costs a fallback. With the pack, the Pinewood agent asked the simulated game after 5 s of game time, still in its menus, got empty answers and grepped the dump anyway. The pack and the game's --help must say when to ask.
  • Bench hygiene: two Arena runs (643k, 253k) saw the answer in our own docs and aren't counted. Examples never use a bench game's names, and this file stays out of bench copies.

What it doesn't measure yet: the simulated games are kinder than real ones. A try takes a second instead of minutes; the dump is about 5,000 lines instead of hundreds of thousands; the game's README says what's on screen. There is one run per task, so no spread yet. Making the bench harder where real games are hard comes first in the pack's probe-and-refine pass.

What it looks like for a developer

An illustration, not a measured run.

"Add Valheim to Fusion."

  1. The agent reads SKILL.md (cached), then runs fusion router recon valheim.exe: Unity Mono, kit unity fits, no anti-cheat, build hash, no existing router.
  2. With Valheim running, it runs fusion router new valheim --auto. That writes a draft router.toml: World, player and camera from Unity defaults, characters found by type.
  3. fusion router check reports: "World: OK. NPCs: 0 found. Valheim's creatures aren't CharacterControllers; try fusion game find Character."
  4. fusion game find Character returns Valheim's own Character class, its live instances and its health getter. The agent adds two lines of data and checks again: everything passes.
  5. The human playtests. The agent writes what it learned to the router's notes.

Today the same job means cloning SkyCraft and writing a BepInEx plugin, a protocol, compositing and collision sampling.

Sources