AnyModel vs OpenCode — local coding agent benchmark

Historical evidence · interpretation clarified October 2026. This is one local Qwen3-Coder 30B model, six small tasks, three repetitions per arm, measured in June 2026. The fixed arm has 17/18 passing artifacts but 16/18 clean completed runs: one passing artifact was produced by a run that timed out. Artifact success is not successful agent completion. This is not a frontier-model comparison, a current hosted-model ranking, or a general speed guarantee. The raw results and original measurements below are retained unchanged. Current product scope · Baseline raw JSON · Fixed raw JSON.

Model: qwen3-coder:30b on Ollama (M1 Max 32GB, 100% GPU) · Wed, 10 Jun 2026 21:17:50 GMT · 54 runs · identical seeded tasks, artifact-verified, fresh workspace per run

AnyModel v1.16.2

Claude Code → anymodel proxy → Ollama

0%

task success (0/18 runs)

  • median turn-around: 3.7s
  • tool calls executed: 9 (bash 3 · edit 0 · read 4)
  • timeouts: 0 · zero-tool turns: 9
  • skill invocations observed: 2

OpenCode v1.16.2

OpenCode → Ollama direct

78%

task success (14/18 runs)

  • median turn-around: 118.4s
  • tool calls executed: 17 (bash 0 · edit 17 · read 0)
  • timeouts: 0 · zero-tool turns: 3
  • skill invocations observed: 0

AnyModel (fixed) v1.17.0

Claude Code → anymodel proxy → Ollama

94%

task success (17/18 runs)

  • median turn-around: 10.2s
  • tool calls executed: 55 (bash 11 · edit 18 · read 10)
  • timeouts: 1 · zero-tool turns: 0
  • skill invocations observed: 6

Median wall time per task (seconds — lower is better)

AnyModel
3.7
OpenCode
118.4
AnyModel (fixed)
10.2

Per-task results (pass / runs · median time · tool calls)

taskAnyModelOpenCodeAnyModel (fixed)
write-file0/33.0s · 0 calls2/3105.4s · 2 calls3/34.3s · 3 calls
bash-count0/33.3s · 1 calls2/3112.0s · 2 calls3/38.6s · 5 calls
fix-bug0/33.8s · 3 calls3/3130.6s · 3 calls3/3283.7s · 29 calls
multi-file0/33.7s · 0 calls3/3119.6s · 6 calls3/311.3s · 6 calls
edit-json0/33.6s · 3 calls3/3120.4s · 3 calls3/310.3s · 6 calls
skill-stamp0/34.8s · 2 calls1/3101.1s · 1 calls2/37.2s · 6 calls

Failures

armrunclasstimeworkspace (logs)
AnyModelwrite-file #1zero-tool-execution4.4s/tmp/bench-runs/baseline-1.16.2/anymodel/write-file/rep1
AnyModelwrite-file #2zero-tool-execution3.0s/tmp/bench-runs/baseline-1.16.2/anymodel/write-file/rep2
AnyModelwrite-file #3zero-tool-execution2.9s/tmp/bench-runs/baseline-1.16.2/anymodel/write-file/rep3
OpenCodewrite-file #3zero-tool-execution103.5s/tmp/bench-runs/baseline-1.16.2/opencode/write-file/rep3
AnyModelbash-count #1zero-tool-execution8.0s/tmp/bench-runs/baseline-1.16.2/anymodel/bash-count/rep1
AnyModelbash-count #2wrong-artifact3.3s/tmp/bench-runs/baseline-1.16.2/anymodel/bash-count/rep2
AnyModelbash-count #3zero-tool-execution2.9s/tmp/bench-runs/baseline-1.16.2/anymodel/bash-count/rep3
OpenCodebash-count #1wrong-artifact111.5s/tmp/bench-runs/baseline-1.16.2/opencode/bash-count/rep1
AnyModelfix-bug #1wrong-artifact5.6s/tmp/bench-runs/baseline-1.16.2/anymodel/fix-bug/rep1
AnyModelfix-bug #2wrong-artifact3.8s/tmp/bench-runs/baseline-1.16.2/anymodel/fix-bug/rep2
AnyModelfix-bug #3wrong-artifact3.2s/tmp/bench-runs/baseline-1.16.2/anymodel/fix-bug/rep3
AnyModelmulti-file #1zero-tool-execution6.8s/tmp/bench-runs/baseline-1.16.2/anymodel/multi-file/rep1
AnyModelmulti-file #2zero-tool-execution3.7s/tmp/bench-runs/baseline-1.16.2/anymodel/multi-file/rep2
AnyModelmulti-file #3zero-tool-execution3.7s/tmp/bench-runs/baseline-1.16.2/anymodel/multi-file/rep3
AnyModeledit-json #1wrong-artifact4.9s/tmp/bench-runs/baseline-1.16.2/anymodel/edit-json/rep1
AnyModeledit-json #2wrong-artifact3.6s/tmp/bench-runs/baseline-1.16.2/anymodel/edit-json/rep2
AnyModeledit-json #3wrong-artifact3.5s/tmp/bench-runs/baseline-1.16.2/anymodel/edit-json/rep3
AnyModelskill-stamp #1zero-tool-execution4.8s/tmp/bench-runs/baseline-1.16.2/anymodel/skill-stamp/rep1
AnyModelskill-stamp #2wrong-artifact5.5s/tmp/bench-runs/baseline-1.16.2/anymodel/skill-stamp/rep2
AnyModelskill-stamp #3wrong-artifact3.7s/tmp/bench-runs/baseline-1.16.2/anymodel/skill-stamp/rep3
OpenCodeskill-stamp #1zero-tool-execution101.0s/tmp/bench-runs/baseline-1.16.2/opencode/skill-stamp/rep1
OpenCodeskill-stamp #3zero-tool-execution101.1s/tmp/bench-runs/baseline-1.16.2/opencode/skill-stamp/rep3
AnyModel (fixed)skill-stamp #1wrong-artifact9.8s/tmp/bench-runs/fixed-1.17.0/anymodel/skill-stamp/rep1
Methodology. Each run: fresh temp workspace, seeded fixtures, one prompt, 300s timeout. Success = artifact verifier only (file contents / passing test), never model self-report. AnyModel arm: claude -p → anymodel proxy → Ollama native API; tool calls counted from Claude Code stream-json events; proxy log sliced per run. OpenCode arm: opencode run → Ollama /v1; tool calls parsed from its execution trace. Same Ollama server, same context length (32768), flash attention + q8_0 KV cache. Skill task seeds the same SKILL.md into .claude/skills, .opencode/skill, and .opencode/skills.