The self-evolving agent framework detailed on arXiv starts from a brutal measurement called the consistency gap.
On AppWorld with a ReAct agent using GPT-4.1, average per-run success is 77% but only 53% of tasks succeed in all five runs.
That 24-point shortfall is not noise, it is the deployment risk.
The authors, Evelyn Duesterwald, Benjamin Elder, Lilian Ngweta, Shashanka Ubaru and Malgorzata Zimon, do not propose a new model.
They propose memory.
