AstraCode 团队实测 11 个模型后发现:每步重新发送上下文导致 step 数成倍放大成本,便宜模型可能因步数多反而更贵;路由优化效果有限。
We build AstraCode, an AI code editor whose agent has to prove its own work. Model calls are the largest line on our bill, so in September we stopped guessing and measured: eleven models, the same set of real coding tasks, each run more than once, scored on pass rate and on cost per finished task.
The cheapest model per token is not the cheapest model per task. On one task a model took 130 agent steps; another finished in 32. Every step re-sends context, so the step count multiplies everything else.
It is not as simple as "expensive models wander more" either. Some of the priciest models took fewer steps than the cheap one. We wrote that rule down, then had to cross it out.
Takeaway: log steps per task alongside tokens. A model that is 3x cheaper per token and takes 4x the steps is a more expensive model.
The obvious optimisation is a router: read the request, send easy ones to a small model and hard ones to a big one. We tried three versions of it.
A per-step router came out 0.4% worse than not routing at all.
A classifier that predicted difficulty from the prompt reached a rank correlation of about 0.75 and still lost.
Even an oracle with perfect hindsight picked the small model on every task, so there was no headroom to win.
The reason is variance. The same small model passed one task 67% of the time on the identical prompt. Whether a run succeeds is decided during the attempt, not by the prompt, so no prompt classifier can see it.
Check its work. Run the tests, then undo the change and confirm those tests fail without it, so a test that passes either way doesn't count as proof.
Only if a check fails, hand the task to a stronger model.
Pass rate went from 0.64 to about 1.00, at roughly $0.20 per task against $0.03 for the small model alone.
These are our tasks and our harness, so treat them as one data point, not a ranking. Every model we tested is good at something.
AstraCode no longer asks you to pick a model. AstraOne starts with the cheaper model and escalates when a check fails. The checks are the same ones you'd want anyway: tests that are shown to fail without the change, a diff you review hunk by hunk, and a checkpoint before every turn.
If you're building on LLMs, two habits are worth more than any model choice:
AstraCode is free to start, no card: https://astracode.io