Claude Code → anymodel proxy → Ollama
task success (0/18 runs)
OpenCode → Ollama direct
task success (14/18 runs)
Claude Code → anymodel proxy → Ollama
task success (17/18 runs)
| task | AnyModel | OpenCode | AnyModel (fixed) |
|---|---|---|---|
| write-file | 0/33.0s · 0 calls | 2/3105.4s · 2 calls | 3/34.3s · 3 calls |
| bash-count | 0/33.3s · 1 calls | 2/3112.0s · 2 calls | 3/38.6s · 5 calls |
| fix-bug | 0/33.8s · 3 calls | 3/3130.6s · 3 calls | 3/3283.7s · 29 calls |
| multi-file | 0/33.7s · 0 calls | 3/3119.6s · 6 calls | 3/311.3s · 6 calls |
| edit-json | 0/33.6s · 3 calls | 3/3120.4s · 3 calls | 3/310.3s · 6 calls |
| skill-stamp | 0/34.8s · 2 calls | 1/3101.1s · 1 calls | 2/37.2s · 6 calls |
| arm | run | class | time | workspace (logs) |
|---|---|---|---|---|
| AnyModel | write-file #1 | zero-tool-execution | 4.4s | /tmp/bench-runs/baseline-1.16.2/anymodel/write-file/rep1 |
| AnyModel | write-file #2 | zero-tool-execution | 3.0s | /tmp/bench-runs/baseline-1.16.2/anymodel/write-file/rep2 |
| AnyModel | write-file #3 | zero-tool-execution | 2.9s | /tmp/bench-runs/baseline-1.16.2/anymodel/write-file/rep3 |
| OpenCode | write-file #3 | zero-tool-execution | 103.5s | /tmp/bench-runs/baseline-1.16.2/opencode/write-file/rep3 |
| AnyModel | bash-count #1 | zero-tool-execution | 8.0s | /tmp/bench-runs/baseline-1.16.2/anymodel/bash-count/rep1 |
| AnyModel | bash-count #2 | wrong-artifact | 3.3s | /tmp/bench-runs/baseline-1.16.2/anymodel/bash-count/rep2 |
| AnyModel | bash-count #3 | zero-tool-execution | 2.9s | /tmp/bench-runs/baseline-1.16.2/anymodel/bash-count/rep3 |
| OpenCode | bash-count #1 | wrong-artifact | 111.5s | /tmp/bench-runs/baseline-1.16.2/opencode/bash-count/rep1 |
| AnyModel | fix-bug #1 | wrong-artifact | 5.6s | /tmp/bench-runs/baseline-1.16.2/anymodel/fix-bug/rep1 |
| AnyModel | fix-bug #2 | wrong-artifact | 3.8s | /tmp/bench-runs/baseline-1.16.2/anymodel/fix-bug/rep2 |
| AnyModel | fix-bug #3 | wrong-artifact | 3.2s | /tmp/bench-runs/baseline-1.16.2/anymodel/fix-bug/rep3 |
| AnyModel | multi-file #1 | zero-tool-execution | 6.8s | /tmp/bench-runs/baseline-1.16.2/anymodel/multi-file/rep1 |
| AnyModel | multi-file #2 | zero-tool-execution | 3.7s | /tmp/bench-runs/baseline-1.16.2/anymodel/multi-file/rep2 |
| AnyModel | multi-file #3 | zero-tool-execution | 3.7s | /tmp/bench-runs/baseline-1.16.2/anymodel/multi-file/rep3 |
| AnyModel | edit-json #1 | wrong-artifact | 4.9s | /tmp/bench-runs/baseline-1.16.2/anymodel/edit-json/rep1 |
| AnyModel | edit-json #2 | wrong-artifact | 3.6s | /tmp/bench-runs/baseline-1.16.2/anymodel/edit-json/rep2 |
| AnyModel | edit-json #3 | wrong-artifact | 3.5s | /tmp/bench-runs/baseline-1.16.2/anymodel/edit-json/rep3 |
| AnyModel | skill-stamp #1 | zero-tool-execution | 4.8s | /tmp/bench-runs/baseline-1.16.2/anymodel/skill-stamp/rep1 |
| AnyModel | skill-stamp #2 | wrong-artifact | 5.5s | /tmp/bench-runs/baseline-1.16.2/anymodel/skill-stamp/rep2 |
| AnyModel | skill-stamp #3 | wrong-artifact | 3.7s | /tmp/bench-runs/baseline-1.16.2/anymodel/skill-stamp/rep3 |
| OpenCode | skill-stamp #1 | zero-tool-execution | 101.0s | /tmp/bench-runs/baseline-1.16.2/opencode/skill-stamp/rep1 |
| OpenCode | skill-stamp #3 | zero-tool-execution | 101.1s | /tmp/bench-runs/baseline-1.16.2/opencode/skill-stamp/rep3 |
| AnyModel (fixed) | skill-stamp #1 | wrong-artifact | 9.8s | /tmp/bench-runs/fixed-1.17.0/anymodel/skill-stamp/rep1 |
claude -p → anymodel proxy → Ollama native API; tool calls counted from Claude Code stream-json events; proxy log sliced per run. OpenCode arm: opencode run → Ollama /v1; tool calls parsed from its execution trace. Same Ollama server, same context length (32768), flash attention + q8_0 KV cache. Skill task seeds the same SKILL.md into .claude/skills, .opencode/skill, and .opencode/skills.