Main results.
| Success Rate (%) | Basic Task | Advanced Task | |||
|---|---|---|---|---|---|
| Harness | Model | Zero-shot | Zero-shot | In-Context Learning |
Skill Learning |
| Hermes | Qwen3.8-Flash | 58.9 | 8.0 | 44.0 | 33.0 |
| GPT5.6-Terra | 53.0 | 11.0 | 16.0 | 34.0 | |
| OpenClaw | Qwen3.8-Flash | 50.6 | 5.0 | 35.0 | 38.0 |
| GPT5.6-Terra | 47.6 | 12.0 | 39.0 | 22.0 | |
| Codex | Qwen3.8-Flash | 52.4 | 4.0 | 34.0 | 33.0 |
| GPT5.6-Terra | 54.8 | 11.0 | 52.0 | 26.0 | |
| Claude Code | Qwen3.8-Flash | 42.9 | 3.0 | 27.0 | 31.0 |
| GPT5.6-Terra | 51.8 | 9.0 | 31.0 | 25.0 | |
(1) Zero shot: the agent attempts tasks directly without prior experience. (2) In-Context Learning (ICL): the agent first completes two related basic tasks, retaining the full interaction history in context for the advanced task. (3) Skill Learning: the same learning phase as ICL, but the agent distills its experience into a concise skill summary, which replaces the raw trajectories at test time.