Your Agent Aced the Task. Will It Do It Again?
Your Agent Aced the Task. Will It Do It Again? Enterprise Article Published September 15, 2026 Upvote 112 Evelyn Duesterwald evduester ibm-research Lilian Ngweta lilianngweta ibm-research Vatche Isahagian Vatche ibm-research Jayaram Radhakrishnan jayaramkr ibm-research Vinod Muthusamy vinodmut ibm-research Gaodan Fang gaodan-fang ibm-research Ashwath Vaithinathan Aravindan ashwath-vaithina ibm-research Punleuk Oum illeatmyhat ibm-research G Thomas gsthomasx ibm-research Merve Unuvar mrvnvr ibm-research Ayhan Sebin ayhansebin ibm-research Michał Ulewicz Michal ibm-research Your agent works in rehearsal, but during the live demo, it takes a different path and fails the same task. In production, it is a reliability problem: a workflow that succeeded once may fail the next time a user makes the same request.
This Research is relevant to the technology intelligence record because it involves GitHub, Intel, gpt-4.1, gpt-oss. The source article should remain the factual reference for follow-up coverage.
- Enterprise Article Published September 15, 2026 Upvote 112 Evelyn Duesterwald evduester ibm-research Lilian Ngweta lilianngweta ibm-research Vatche Isahagian Vatche ibm-research Jayaram Radhakrishnan jayaramkr ibm-research Vinod Muthusamy vinodmut ibm-research Gaodan Fang gaodan-fang ibm-research Ashwath Vaithinathan Aravindan ashwath-vaithina ibm-research Punleuk Oum illeatmyhat ibm-research G Thomas gsthomasx ibm-research Merve Unuvar mrvnvr ibm-research Ayhan Sebin ayhansebin ibm-research Michał Ulewicz Michal ibm-research Your agent works in rehearsal, but during the live demo, it takes a different path and fails the same task.
- In production, it is a reliability problem: a workflow that succeeded once may fail the next time a user makes the same request.
- For mission-critical work, such as reconciling a financial transaction or checking a contract for an obligation, that can be a showstopper.
- Most benchmarks hide this variability behind an average.
- On AppWorld, a ReAct agent using GPT-4.1 succeeded on 77.4% of runs across five repetitions.
- But it succeeded in all five runs for only 53.0% of tasks — a 24.4-point consistency gap .