EmbodiedCLBench:Evaluating Continual Learning
for Self-Evolving Embodied Agents

The first benchmark in embodied simulation for evaluating the continual learning capability of self-evolving agents

Jie Huang★,1,2 Yanan Chen★,1 Ruixun Liu★,1 Kaichen He1
Xinyi He1 Rui Huang3 Zilong Zheng2 Zhenliang Zhang2 Yiwu Zhong1,2,†
1School of Intelligence Science and Technology, Peking University 2State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing, China 3University of Hong Kong
★ Equal contribution. † Corresponding author.
EmbodiedCLBench overview showing compositional generalization from basic to advanced tasks
Core design: we operationalize continual learning through compositional generalization: an agent first completes simpler tasks to accumulate experience, then is tested on whether it can recombine the acquired skills to solve more complex, unseen tasks.

Abstract

Continual learning requires models to accumulate past experience and transfer it to new tasks. Despite substantial research on continual learning, existing work remains largely confined to non-embodied settings, without extending to embodied environments.

To address this gap, we introduce EmbodiedCLBench, the first benchmark in embodied simulation for evaluating the continual learning capability of self-evolving agents. With compositional generalization as the core evaluation principle, our benchmark evaluates whether agents can first learn from basic tasks and then recombine the acquired skills to solve complex, unseen advanced tasks. Based on EmbodiedCLBench, we conduct extensive experiments across multiple agent harnesses and learning methods, revealing consistent limitations and distinctive behaviors.

Our results show that while advanced tasks can be hardly solved under zero-shot setting, the experience from basic tasks enables consistent improvement. Such benefit from the experience remains stable even if the task complexity increases. However, this improvement does not scale accordingly with the number of learning samples, and current agents cannot extract meaningful information from additional experience. Collectively, our findings highlight continual learning as a fundamental yet underexplored capability, and our benchmark offers valuable resources for designing effective self-evolving agents in embodied environments.

Benchmark Construction

Overview of the EmbodiedCLBench construction process
Overview of the construction process
EmbodiedCLBench benchmark statistics
Benchmark statistics

Results and Findings

Main results.

Success Rate (%) Basic Task Advanced Task
Harness Model Zero-shot Zero-shot In-Context
Learning
Skill
Learning
Hermes Qwen3.8-Flash58.98.044.033.0
GPT5.6-Terra53.011.016.034.0
OpenClaw Qwen3.8-Flash50.65.035.038.0
GPT5.6-Terra47.612.039.022.0
Codex Qwen3.8-Flash52.44.034.033.0
GPT5.6-Terra54.811.052.026.0
Claude Code Qwen3.8-Flash42.93.027.031.0
GPT5.6-Terra51.89.031.025.0

(1) Zero shot: the agent attempts tasks directly without prior experience. (2) In-Context Learning (ICL): the agent first completes two related basic tasks, retaining the full interaction history in context for the advanced task. (3) Skill Learning: the same learning phase as ICL, but the agent distills its experience into a concise skill summary, which replaces the raw trajectories at test time.

Effect of Compositional Complexity

Does more complex composition make continual learning harder?

Effect of compositional complexity on continual learning

Scaling with the Number of Learning Samples

Does more experience lead to a higher success rate on advanced tasks?

Success rate scaling with the number of learning samples

BibTeX

@misc{huang2026embodiedclbench,
  title   = {EmbodiedCLBench: Evaluating Continual Learning
             for Self-Evolving Embodied Agents},
  author  = {Huang, Jie and Chen, Yanan and Liu, Ruixun and
             He, Kaichen and He, Xinyi and Huang, Rui and
             Zheng, Zilong and Zhang, Zhenliang and Zhong, Yiwu},
  year    = {2026},
  url     = {https://pku-value-lab.github.io/EmbodiedCLBench-homepage/},
}