Average success is the wrong number if you plan to leave an agent running.
Evelyn Duesterwald, Benjamin Elder, Lilian Ngweta, Shashanka Ubaru and Malgorzata Zimon, writing from IBM Software Innovation Lab and IBM Research, put a name on the hole. On the 168-task test-normal split of AppWorld, a ReAct agent on GPT-4.1 posts a 77% mean pass rate across five runs. It clears all five on only 53% of tasks. They call the 24-point difference the consistency gap, and they treat it as a deployment metric, not a lab curiosity.
About a third of the benchmark is mixed: the same task sometimes works and sometimes does not. That is the slice where a published average is least like what a user sees.
Difficulty makes it worse for the stronger model. On GPT-4.1 the gap is 17.5 points on easy tasks, 25 on medium, 30.2 on hard. A weaker open model, GPT-OSS-120B, looks different and worse: 34% mean, 10% all-five, and zero all-five wins on hard tasks it can occasionally solve. Normalized consistency (all-five divided by the mean) is 0.69 for GPT-4.1 and 0.30 for the open model.
They also argue you cannot decode your way out. Flat next-token distributions flip under the small noise of a hosted endpoint, even at temperature zero. Greedy decoding and a fixed seed only freeze how a distribution is turned into a token. They do not stop the probabilities themselves from jittering.
