AI Agent Aiden Outperforms Humans in OpenAI Challenge

Weco's AI agent, Aiden, topped OpenAI's 'Parameter Golf' challenge, outperforming human engineers and demonstrating the power of AI in research and development.

5 min read
Presentation slide showing a graph with bars representing leaderboard contributions, with 'dexhunter (Aiden)' clearly leading.
AI Engineer
Visual TL;DR
OpenAI Parameter GolfDriver
challenge to train best language models under size and computation constraints
From the articleIn a significant development for the field of artificial intelligence and machine learning, an AI agent named Aiden has emerged as the top contributor in OpenAI's recent hiring challenge, 'Parameter Golf'.
Weco's AI Agent AidenCore
From the article 3 mentionsZhengyao Jiang, co-founder and CEO of Weco, presented Aiden's remarkable achievement, explaining that the agent is a self-improving system capable of processing public information like research papers and pull requests, running experiments, and submitting its findings.
Aiden Outperforms HumansOutcome
From the article 4 mentionsAiden, developed by the auto-research company Weco, outperformed human engineers by setting seven leaderboard records, more than twice the number achieved by the best human participant.
AI in R&DEffect
demonstrating the power of AI in research and development tasks
Community ImpactContext
Aiden's success sparked discussions on AI's role and future collaboration
From the article 2 mentionsJiang emphasized that Aiden's success wasn't solely measured by scores but also by its impact within the human community.
Human-AI SynergyContext
highlights the synergistic nature of human-AI collaboration in complex tasks
Evolving AI EngineeringEffect
shifts focus from manual tuning to designing self-improving AI systems
Contents(5)

In a significant development for the field of artificial intelligence and machine learning, an AI agent named Aiden has emerged as the top contributor in OpenAI's recent hiring challenge, 'Parameter Golf'. The competition, which aimed to train the best language models under size and computation constraints, saw approximately 1,000 participants and over 2,000 submissions. Aiden, developed by the auto-research company Weco, outperformed human engineers by setting seven leaderboard records, more than twice the number achieved by the best human participant.

AI Agent Aiden Outperforms Humans in OpenAI Challenge - AI Engineer
AI Agent Aiden Outperforms Humans in OpenAI Challenge, from AI Engineer

Aiden's Performance in Parameter Golf

Zhengyao Jiang, co-founder and CEO of Weco, presented Aiden's remarkable achievement, explaining that the agent is a self-improving system capable of processing public information like research papers and pull requests, running experiments, and submitting its findings. During the 22-day competition, Aiden ran approximately 1,300 experiments on a single H100 node, utilizing only 4% of the total compute available. Aiden's submissions had a 28% acceptance rate, six times higher than the community average, significantly boosting the signal-to-noise ratio within the competition's communication channels.

Beyond Scores: Community Impact

Jiang emphasized that Aiden's success wasn't solely measured by scores but also by its impact within the human community. The agent's work was forked, cited, and built upon by other participants, evidenced by its PR-citation h-index of 10, compared to the next highest human's score of seven. This demonstrates the AI's ability to produce work that is not only high-performing but also recognized and valuable to human collaborators.

The Synergistic Nature of Human-AI Collaboration

Delving into the origins of Aiden's success, Jiang highlighted that the AI's record-breaking contributions were largely derived from existing human research. Aiden was adept at identifying and implementing ideas from research papers, other participants' attempts, and even notes from human researchers who had abandoned ideas due to implementation difficulties. A small fraction of Aiden's original ideas emerged from its efforts to navigate file size constraints. One notable example involved Aiden combining an idea from a paper on 'gated attention' with a subsequent tokenizer improvement posted by another contributor, leading to a significant performance leap.

The Craft of Autoresearch

Jiang drew a parallel between doing autoresearch and training a model, stating that the code base abstraction acts as the architecture, setting constraints and priorities, while the evaluation serves as the loss function and data. He stressed the importance of designing effective evaluation metrics and robust code base abstractions, as these elements guide the AI's search and learning process. Jiang illustrated this with an example from a fraud detection pipeline, where a loose API led to data leakage, while a tightened abstraction resolved the issue.

The Evolving Role of AI Engineers

Jiang concluded by framing autoresearch as a new craft, where designing effective evaluations and abstractions is paramount. He predicted that creativity and judgment in these areas will become increasingly valuable for AI engineers. The rise of autoresearch, he suggested, is akin to how gradient descent revolutionized coding, commoditizing certain execution skills while elevating higher-level strategic thinking. This shift means that humans will move up the stack, guiding AI systems rather than being replaced by them.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.