General Reasoning Founders on Scaling AI Models to Long Horizons
Ross and Chengxi Taylor of General Reasoning discuss Galactica, RLHF, KellyBench, and the infrastructure needed to scale AI models to long horizons.

Visual TL;DR
From the article 6 mentionsAt the AI Engineer World's Fair, Ross Taylor and Chengxi Taylor, co-founders of London-based research lab General Reasoning, delivered a joint keynote on expanding machine intelligence beyond short text generations.
initial reasoning recipes struggled to scale effectively for complex, extended tasks
today's frontier models fail at real-world strategy, highlighting current limitations
From the articleTo evaluate current frontier models on extended real-world strategy, General Reasoning created KellyBench.
infrastructure and ecosystem limitations hinder the development of advanced AI models
From the article 4 mentionsChengxi Taylor is the co-founder and president of General Reasoning, focusing on agent environments, credit assignment, and compute optimization for extended reasoning trajectories.
expanding machine intelligence beyond short text generations to complex, multi-week objectives
From the article 9+ mentionsDrawing from their past research on Meta's Galactica and Llama models, the pair mapped out the algorithmic, environment, and compute infrastructure necessary for scaling AI agents to complex, multi-week objectives.
base models alone are insufficient for long-horizon reasoning, needing more than just text
From the article 6 mentionsTaylor noted that the missing ingredient was a pure application of the bitter lesson.
algorithmic, environment, and compute infrastructure needed for scaling AI agents
From the articleTaking over the presentation, Chengxi Taylor defined long horizon tasks not merely as engineering hurdles, but as a long-term mindset required to tackle civilization-scale problems like drug discovery or mathematical proofs.
developing robust AI agents capable of complex, multi-week objectives and real-world strategy
From the article 5 mentionsWithout larger context windows, stronger base models, and massive reinforcement learning compute scaling, agents could not develop emergent backtracking mechanisms.
Contents(6)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.