Give every task a role.
Match research, planning, building and review to the tools and model tiers in your own stack.
How routing works ↗MODEL ROUTING FOR AI CODING AGENTS
Give your AI a playbook for choosing the model, effort, and tools each task needs. Routing rules, subagents, and a CLI runner, built around your setup.
npx model-orchestratorRole picks the job. Complexity sets the effort.
Your installed stack supplies the available models.
A quick look at model-orchestrator. Silent video.
Match research, planning, building and review to the tools and model tiers in your own stack.
How routing works ↗Send each worker a task brief, shared facts, and clear acceptance checks through the lane runner.
Meet the aunx commands ↓Use acceptance checks and independent review. Measure where your agent sends work with local routing metrics.
See the evidence ↓Choose your setup and the AI tools you use.
Confirm the project rules and activation steps. Existing files get backups.
Follow the activation check in your generated README, then give your agent a real task.
Want to inspect the plan first? Use --dry.
Read the complete installation guide ↗
Dated measurements, sample sizes, and scripts you can rerun. The guarantees guide distinguishes executable checks from instructions your agent follows.
Latest published version, checked with npm at build time.
A refused symlinked or changing file (such as MANIFEST.json) now names its path and the fix: replace the link with the file it points to. Since 1.0.1, rerunning the installer over a symlinked manifest stops instead of treating it as a fresh install.
A path shown in a refusal escapes control characters, so a crafted file name cannot write terminal escape sequences; a directory or other non-file gets "expected a regular file" without symlink advice.
aunx route-metrics always runs the packaged metrics implementation. A checkout can no longer substitute its own JavaScript when a user requests a summary.
Installer and uninstall manifests use bounded regular-file reads with symlink refusal. Acceptance-check and routing JSON reads share the same bounded reader.
The aunx command. The existing package now ships a short command alongside model-orchestrator. Run the installer with aunx, call an AI through aunx cli-run, scaffold a task brief and context file, execute acceptance checks, get a routing suggestion or summarize your local routing activity.
A build process from requirements to verified use. Shared context, executable acceptance checks, live capability probes, scoped assignments, background heartbeat guidance, split-build ownership and one independent audit step with a companion reviewer.
The description names all three things the installer writes. npm, the GitHub About box and package.json now read: rules, subagents and a lane runner tell agents which model to use. The README opening already listed the lane runner; the description stopped at rules and subagents.
Every GitHub topic is an npm keyword. coding-agents was a topic with no matching keyword.
Opt-in snippet application. --apply-snippets applies a replaceable Claude Code rules block and merges hooks with timestamped backups, validation before writes, dry-run previews and manual uninstall guidance. (#41)
Model orchestrator or a model proxy. README and llms.txt explain the request layer, when to pick each approach and how they compose. (#44)
docs/install.md.Manifest-based uninstall. --uninstall --dir <dir> --project <project> removes managed files whose content matches the recorded hash, keeps and names edited files, and preserves files outside the manifest. --dry and --dry-run preview the same removals. The manifest is the last file removed and stays when edits remain. New installs record the directories they create so uninstall can remove them when empty; older manifests leave directories in place. The command prints the manual steps for removing pasted rules and merged hooks. A missing manifest exits 2 and names its expected path.
Activation and related tools in the README. A captured install supplies the rules, hooks and smoke-test steps. A shared table links agent-personalizer and website-build-skill, and a short uninstall section links the full instructions.
CLAUDE.snippet.md said "Two hooks were written" and named only route-gate.mjs and subagent-context.mjs, while a claude-code install writes and wires route-metrics.mjs too. It now names all three and how to read the metrics log. The same pass corrects llms.txt and docs/install.md (three hooks on a full install), templates/agents/README.md (lists route-metrics.mjs), and the README plugin section, which said the plugin ships three hooks: it ships the two read-only ones, and route-metrics comes only with the npm install. A new test fails if the snippet names fewer hooks than the plan writes.A recording of the command doing its job, at the top of the README. docs/demo.gif (60 KB) shows npx model-orchestrator ... --dry typed and run: the level, the AIs, both target folders, all 38 files it would write, and the closing line that nothing was written. scripts/record-demo.sh re-records it from the PUBLISHED package inside a temp folder, so the frames stay the program's own output rather than a staged screen, and anyone can reproduce them with brew install asciinema agg. The text plan stays in the README under the image, and the image carries alt text describing what it prints, so the page still reads with images off.
--summary now answers the question the package exists for: how much work left the main session. work sent off the main session counts covered turns whose route marker named any lane that is not an inline name, over covered turns. It is derived only from lane names the log already holds, with no price table and no token estimate, because the log holds neither. --inline a,b renames what counts as inline, since the lane vocabulary belongs to your own ROUTING.md. A log with no lane markers yet says so instead of printing a number.
--dry plan. The old opening was a 200-word paragraph followed by "What it is not", the flag-conflict rules and the uninstall procedure, all before the reader had seen the tool do anything. The detail moved rather than went away: docs/install.md (every flag, the two target folders, headless examples, the full file list), docs/how-it-routes.md (role, complexity and stakes; the three verifier agents; pinning model and effort), docs/guarantees.md (enforced by code, delegated to a vendor flag, or only an instruction), docs/companions.md (the full companion-tool table). docs/README.md and llms.txt index all four. Platform support and the vendor compatibility table stay in the README inside collapsed <details> blocks, because scripts/gen-catalog.js writes the vendor table between markers there and test/prose.test.js requires the skipped-test explanations to live in the README; collapsing them keeps both mechanisms pointed at the same file. The opening line still satisfies test/copy.test.js: it names the package, carries the shared purpose clause, and keeps the 0.1.11 correction that it does not select models itself.cmd /c snippet, and OBSIDIAN-TC.md now says so with the code it was read from. All five obsidian-tc snippets spawn "command": "npx", and on Windows npx is npx.cmd, which a bare CreateProcess will not start. Each client was read instead of assumed, and all five resolve it themselves: Zed hands the command to the system shell (crates/context_server/src/transport/stdio_transport.rs, ShellBuilder::new(&Shell::System, ..), from its PR #42382, 2025-12-10), VS Code's formatSubprocessArguments in src/vs/workbench/api/node/extHostMcpNode.ts resolves the executable and re-spawns with shell: true for a .bat or .cmd, the Codex CLI calls which::which_in in codex-rs/rmcp-client/src/program_resolver.rs, and Cursor and Antigravity's agy both reach npx through @modelcontextprotocol/sdk, whose client/stdio.js imports cross-spawn and re-invokes a non-.exe as %COMSPEC% /d /s /c. That last point is wider than those two clients: every published SDK from 1.23.0 through 1.30.0 depends on cross-spawn ^7.0.5, so shell: false in that transport is not the whole story. The doc also names the one Windows case that can still fail and is not about .cmd: Zed prefers PowerShell, which resolves a bare npx to npm's npx.ps1 shim, and a .ps1 will not run under the Restricted execution policy that is Windows' client default. A test pins the snippet set, fails if any .windows file appears, and fails if the doc drops a citation. Not run on a Windows machine.mcp/context7.zed.settings.json spawns "command": "npx", and on Windows npx is npx.cmd, a batch file a client cannot spawn as a bare command. New mcp/context7.zed.windows.settings.json runs "command": "cmd" with "args": ["/c", "npx", "-y", "@upstash/context7-mcp@<pin>"], the shape Context7's own client guide ships for Windows. CONTEXT7.md gains a "Local npx on Windows" table saying which Zed snippet to use on which OS, and the same by-hand change for any other client taking the local alternative. The remote snippets spawn nothing and are unchanged. A test pins the Windows args to the POSIX args behind /c npx and the catalog pin; not yet run on a Windows machine.Context7 as a third companion tool, paired with codecalc. Context7 (Upstash) hands the agent current, version-specific documentation and code examples for any library, SDK, API or CLI, hosted or run locally with npx. It answers what a library is documented to do; codecalc still answers what the code actually does by running it. A new protocol, protocols/docs-then-prove.md, states the rule the pairing serves: pull current docs before writing against anything unconfirmed this session, then a run proves it, and where a doc and a run disagree the run wins. Optional and off by default, like obsidian-tc (--tools context7, or the interactive question); unlike the other two companions it always makes a network call, so it is the one to skip offline. src/catalog.js carries the entry (repo, role, install, requirements, the clients it self-registers with); CONTEXT7_STATUS renders both selected and not-selected wording the same way CODECALC_STATUS and OBSIDIAN_TC_STATUS already do, at every level and in ROUTING.md.
templates/tools/context7/CONTEXT7.md plus seven per-client mcp/ snippets, each read from that client's own docs rather than shared across clients that do not actually share a config shape: context7.claude-code.mcp.json ("type": "http" next to url, required or Claude Code skips the server), context7.mcpServers.json (Cursor's own documented shape, url with no type), context7.vscode.mcp.json, context7.qwen.settings.json (httpUrl, not url, plus the non-credential Accept header upstream ships), context7.zed.settings.json (local npx, pinned to the catalog's context7 version), context7.codex.config.toml, context7.agy.mcp_config.json. Claude Desktop has no file: its remote connection is a UI step (Settings > Connectors > Add Custom Connector), documented in CONTEXT7.md instead of a snippet nothing there reads.
cli-run says WHY a lane failed. Every run lands in one class with its own exit code: auth 14, quota 15, rejected 16, refused 17, cut_short 18, next to the existing empty 10, no_output 11, timeout 12 and unavailable 13. A missing API key, a spent quota and an unknown model id used to share exit 10 or the vendor's own code, and each needs a different response: a missing key is not a model fault, and retrying a spent quota cannot help. Signals are read from each lane's authoritative error fields only, never the model's prose, with precedence auth, quota, rejected, refused, cut_short, empty.
refused=N on every run. Tool calls a hook or deny rule blocked, counted from qwen's permission_denials, agy's deny-rule steps, codex's router Rejected( lines, and grok's session transcript (its sessionId is charset-checked, the path is contained to the sessions root, lines must name the same session, and the read is capped at 5 MiB). null when a lane gives no signal. A deliverable with refused calls is still exit 0.
grok plan (high headroom). xAI's pricing page lists it with "Significantly higher usage across Chat, Imagine, Voice & Build", so --plans grok=supergrok-plus now works and counts as a high-headroom lane for plan guidance and --effort-auto. SuperGrok Heavy stays out: the page states no Build usage for it. A test holds both.Subscription plans, stated by you and never guessed. The installer asks which plan you hold for Claude Code, Codex, agy and grok (interactive, or --plans codex=pro-20x,agy=ultra-5x; --plans none clears). Each plan row in src/catalog.js carries only a name, a headroom level (base, high, max), its official source page and the date it was checked; no prices and no model ids, because both change faster than releases. --list prints them.
Plan guidance in the generated docs. The lanes table gains a Plan column, and ROUTING.md and DELEGATION_MATRIX.md gain a plan guidance block: high and max headroom lanes take volume (scoped builds, pre-ship second-family audits, first-pass research), base headroom lanes keep short second opinions. Headroom moves volume only; who reviews what does not change.
A Claude Code plugin. /plugin marketplace add aunysillyme/model-orchestrator, then /plugin install model-orchestrator@model-orchestrator, installs route-gate.mjs, subagent-context.mjs and the eight subagents without merging a settings snippet by hand. The bundle lives in plugin/, listed by .claude-plugin/marketplace.json at the repo root. It passes claude plugin validate --strict, the check Anthropic's community marketplace review runs, and all eight checks of Sigistry's public plugin verification methodology (1.2), run standalone before release as a quality bar; the plugin is not listed there.
The plugin is generated, never a second copy. npm run gen:plugin renders plugin/ from the same templates/ the installer uses, and test/plugin.test.js fails when the committed bundle drifts from that, when plugin.json's version is not package.json's, when hooks/hooks.json references a hook that is not shipped, when a plugin hook gains a network call, a file write, credential access, dynamic evaluation or a subprocess, when an agent has no tools: line or a review-type agent carries Write or Edit, and when the plugin README loses its install commands. Each check was proved red against the real files before it was trusted.
The route-gate.mjs non-regular-file guard is now tested on every OS, not only where mkfifo exists. The FIFO test is the only one that can prove the HANG the guard exists to prevent (a naive readFileSync on a writer-less FIFO blocks forever), and Windows has no mkfifo to build one, so that test was skipped there and nothing exercised !st.isFile() on Windows at all. A directory at the same rules path reaches the same guard before any open or read call, on every OS, so the guard itself is covered everywhere and the win32 skip is no longer its only coverage.
A guard that refuses an undocumented test skip. test/prose.test.js pins each skip to its file, its exact marker and a phrase the README has to carry, then counts every { skip: in the suite and fails if the totals disagree. Proved in both directions: adding a skip anywhere fails it, and removing a skip's explanation from the README fails it.
The level 3 box's weekly audit could record a hung --version probe as a version string instead of "UNVERIFIED: timed out". weekly-audit.sh's killtree killed a hung process's children before the process itself, so in that gap the probe's own shell could print and exit 0, and bounded() reported success. Seen once on macOS CI ("codex never" in the report); 0 of 20 local runs reproduced it, so it is rare but real on the box. killtree now freezes each process (SIGSTOP) before walking its children, and the watchdog leaves a marker when it fires so bounded() returns 124 whatever order the kills land in. A new test forces the bad ordering and checks the timeout still reads as one.
On Windows, bin/cli-run.mjs could not run any lane at all: spawn() threw EINVAL for every .cmd binary, which is how npm installs every agent CLI there. Since Node's fix for CVE-2024-27980 (18.20.2, 20.12.2, and every 22.x), spawning a .bat/.cmd target without shell: true throws instead of silently running it through an unsafely-escaped cmd.exe. This was a real, shipped defect, not a test gap: a Windows user following this README could not have run a single lane before this release. bin/cli-run.mjs now resolves the .cmd shim to the Node script npm's own cmd-shim tool wrote underneath it and spawns Node directly on that script (resolveCmdShim, windowsSpawnPlan), so a prompt (untrusted text this tool does not control) never passes through a shell at all in the common case. A lane whose .cmd/.bat cannot be resolved that way (an old or hand-edited shim) is refused with exit 13 and a message saying how to fix it, never run through cmd.exe: a batch file re-reads its arguments through %* after cmd.exe has parsed them once, and no escaping fully contains a prompt through both passes. The escaped cmd.exe path (the algorithm documented at qntm.org/cmd and used by cross-spawn, with windowsVerbatimArguments) survives only as an explicit opt-in for the installer's own npm install -g <pinned spec>, whose arguments never include user text. Verified against the real, byte-for-byte output of cmd-shim@9.0.2 (the package npm itself uses), not a guessed shape; the escaping is pinned to exact expected strings for &, |, ^, %, ", a trailing backslash and a literal newline in test/judges.test.js.
README.md and TIERS.md), "attack lane" / "Stage 5 Attack" / "attack pass" are now "challenge lane" / "Stage 5 Challenge" / "challenge pass", "adversarial" (auditor, read, critique, turn, pass) is now "second-opinion", "blast radius" is now "everything it touches", "fail(s) closed" is now "refuses by default", and "threat model" is now "security notes" in the files that link to it. test/prose.test.js gained a permanent check (no alarming security wording in user-facing text) over the purely-prose, user-facing surface (docs/, templates/, README.md, llms.txt, CONTRIBUTING.md, the PR template) so the old wording cannot silently creep back in. docs/audit-brief.md, SECURITY.md, CODE_OF_CONDUCT.md and code identifiers/comments (for example the ATTACK_LANE render var) are unchanged, since these words are expected or load-bearing there.A third claude-code-only hook, route-metrics.mjs, answers "is my agent actually routing and delegating?" A routing rule nobody measures is a rule nobody knows is followed. Wired to five events (UserPromptSubmit, PreToolUse on Agent/Task, SubagentStart, SubagentStop, Stop), it appends one JSON line per event to ~/.ai-orchestrator/route-metrics.jsonl (the same directory and home resolution bin/cli-run.mjs already logs to): a turn, a subagent dispatch (subagent_type, background flag), a subagent start and stop (so a duration can be computed from a small state file keyed by sha256(agent_id)), and the lane parsed from a new hidden marker, <!-- route: <lane> | <why> -->, that the route-gate block now asks every reply to end with. Only named, charset-bounded fields ever reach the log; prompt text, tool descriptions, the raw assistant message, and the marker's "why" half never do. node .claude/hooks/route-metrics.mjs --summary [--since <ISO date>] reports turns, route-marker coverage, lanes by count, dispatches by subagent_type, dispatches with no matching start, and mean/max duration per agent type. Plain Node, zero deps, prints nothing to stdout on any event, fail-open (a miss is a missing log line, never a blocked turn). Installed and wired only when claude-code is the primary, same no-overwrite rules as the other two hooks. See docs/audit-brief.md for the full threat-model writeup.
CI now runs on windows-latest too, node 18/20/22, alongside Ubuntu and macOS. defaults.run.shell: bash makes every workflow step Git Bash on the Windows runner instead of the default pwsh, so the same script runs on all three OSes with no parallel Windows rewrite.
subagentsLoadRules: true on the claude-code catalog entry, with the doc quote as its comment. Drives every new render var below through src/install.js; nothing here is a template branch, per the house rule that templates carry no logic.
Builder executes by default, on claude-code. ROUTING.md rule 5, its "Who builds" section, the "Add an endpoint" example, and the claude-code CLAUDE.snippet.md now say: the orchestrator plans, briefs, verifies and talks to the human; it stays inline only when (a) the brief would cost as much as the work, (b) the task needs this conversation's own context, or (c) it is the human's decision or the final verification of delegated work. Every other primary keeps "the orchestrator builds it directly."
The "Then prove it took" list no longer sends a level 1 reader to a file level 1 never wrote (#27). Step 4 told every reader, at every level, to pick a lane out of bin/lanes.json and run node bin/cli-run.mjs. Level 1 writes no bin/ at all, and step 3 immediately above it hedged correctly with "At level 2+" while step 4 did not. The list is now proofSteps() in src/install.js, gated on level the same way activationSteps() is, and the template renders it. Two tests: the README's section must equal the array exactly for every level and primary, and no bin/ path may appear in it that the plan did not write.
The box setup no longer tells you to sign in to CLIs you did not pick (#26). templates/advanced/vm/README.md step 3 was a fixed sentence naming codex login --device-auth, grok login --device-auth and agy. A level 3 install of claude-code, codex, qwen and ollama was told to sign in to two CLIs it does not have and never told about the one it does. The step now renders each selected CLI's own auth string from the catalog. Everything else in that file was already computed from the selection, which is what made the one hardcoded line easy to miss.
The generated README no longer describes a different install from the one the terminal just printed (#20). The activation list existed twice: once as an array built in bin/cli.js, once as prose in templates/common/README.md that assumed a chat app. A level 2 Claude Code install was told, on the page it was pointed at, to paste PASTE-INTO-YOUR-AGENT.md, a file that run never wrote, and a level 1 chat install was told its rules file was your agent's instructions file, a leftover placeholder. activationSteps() and snippetFor() now live in src/install.js and both surfaces render the same array, so the page can only ever name the file that was written. A test renders every level against every possible primary and fails if the README omits a printed step or names any other agent's snippet.
A chat install no longer claims a project root it never created (#21). Level 1 with a chat app writes no project files, and the README still printed --project as "where your agent reads rules and subagents" next to "subagent definitions: none". It now says there is no project root and why. A CLI primary that reads a rules file but gets no subagent folder (codex, qwen) keeps its project path and gains the missing half: whether this run created that folder.
README.md only. "Route every task to the cheapest AI that does it well" survived in package.json's description, which is what npm search results show, in the repository's GitHub description, which is what GitHub search shows, and in the installer's own banner, printed to every user on every run. All three now say what the package generates instead of what it guarantees: "Routing instructions and a CLI runner for your AI tools." The cheapest-capable-lane guidance in the docs and templates is untouched; that is the product's advice, not a promise about what the code enforces.TimeoutStartSec dies on SIGKILL, so no trap and no cleanup line of ours can run, and its reports/.audit-<stamp>-XXXXXX file was left behind forever. The job now sweeps .audit-* older than a day at start. A day is far outside the unit's own 900s deadline, so a temp belonging to a run still in flight can never be swept. Found by actually starting the unit on Ubuntu; the previous text-only assertion could not see it.--version and -v on the installer, printing the package version and exiting before anything else is validated, so it answers from a broken or half-configured directory.
A "Vendor version compatibility" section in the README, naming the exact vendor CLI version each lane was built against, and saying plainly that the installer checks a binary's presence and never its version. Closes the compatibility-statement item on #11.
author in package.json, so npm shows a byline: aunysillyme (https://github.com/aunysillyme).
The installer's last line now points at the repository, on the reasoning that the end of a successful install is the moment a user is most likely to act on it.
Preserve unrelated subagents in uninstall guidance and explain manual activation cleanup.
Provide a chat activation block below 1,500 characters and explain protocol uploads.
--dry-run as an alias for --dry; the package script was already named dry-run (#16).
cli-run: killTree() ends a lane's process tree with taskkill /T /F on Windows instead of killing only the root process. Windows is still not exercised by CI and stays documented as unsupported; the branch is unit-tested by argv capture (#18).
Published to npm as model-orchestrator (0.1.4 was the first publish, by hand). npx model-orchestrator is now the install line; the GitHub route stays for pinned or unreleased runs.
.github/workflows/release.yml: on a v* tag, checks the tag against package.json, runs the tests, and publishes with provenance through npm trusted publishing (no stored token). Needs the one-time trusted-publisher setup on npmjs.com described in RELEASING.md.
--update-docs: after a selection change, regenerate the documents a previous run wrote and nobody edited since. The check is the same hash rule the runtime class uses: an installed copy that matches the hash MANIFEST.json recorded is regenerated and named under "documents updated"; one that differs is kept and named under "document CONFLICT, kept"; without a manifest every changed document is kept as UNVERIFIABLE. --force still replaces everything; --dry reports and writes nothing. The reconfiguration hint names the flag.Community files: CONTRIBUTING.md, CODE_OF_CONDUCT.md (Contributor Covenant 2.1), MAINTAINERS.md, RELEASING.md, AGENTS.md and CLAUDE.md for contributors' agents, yml issue forms with blank issues disabled, a pull request template, CODEOWNERS, .editorconfig, Dependabot for the workflow actions.
test/prose.test.js: fails on an em dash anywhere in the repo's text files, so the house rule is checked rather than requested.
cli-run: SIGINT/SIGTERM to the wrapper kill the lane's process group before exiting 130/143, with the handlers registered before the spawn so a slow runner cannot signal between the two (#13).
cli-run: stdout and stderr go through streaming UTF-8 decoders and limits are counted in bytes, so a multibyte character split across chunks survives (#14).
cli-run: lanes run in their own process group and the group is killed on timeout or buffer overrun (#1); the durable log stores only a fixed reason code (#4); nonzero vendor exits pass through with an exit_nonzero verdict and a bounded stderr head on the terminal (#9).
Weekly audit: temp-and-rename so a failed rerun never truncates the last good report, failed output kept beside it (#3); every probe under a watchdog, TimeoutStartSec=900, UNVERIFIED lines for timed-out probes (#10).
Installer: three levels, access-aware AI selection, primary-agent loading surface, companion-tool questions (codecalc recommended, obsidian-tc optional), strict flags, containment preflight, exclusive create with rollback, no vendor scripts run.
bin/cli-run.mjs: one entrypoint for grok, codex, agy, hermes and qwen with each lane's native success signal; exit 10 on a run that produced nothing, fail-closed lanes.json, signal handling, digest-only log, --doctor.