Weights & Biases made its agent improve itself

Weights & Biases' Arya logs production traces in Weave and rewrites itself against 886 YAML tasks, now around 66 percent on some families.

Zubin Aysola demoing Arya agent traces in Weave
Live demo of Arya turning production traces into eval tasks· AI Engineer

Zubin Aysola of Weights & Biases demoed Arya, the company's self-improving research agent, live on AI Engineer after its general availability on Monday.

Weights & Biases made its agent improve itself - AI Engineer
Weights & Biases made its agent improve itself, AI Engineer

He followed a main-stage talk by his peer Tim the previous day with a deep dive that reused slides from his NeurIPS presentation.

The core claim is a tight flywheel between production traces and an offline simulation that run bite-wise identical agent code.

On one side Weights & Biases Weave captures every production trace in the same format as offline runs.

On the other side a research sandbox hydrates YAML configs, loads full training logs, and can even simulate GPU executions.

A four-hour sync job mirrors production code into research so hill climbing never drifts.

That loop let Aysola ask Arya to take a live production trace, log it as a task, and spin up training jobs against its own codebase artifact.

The demo prompt instructed the agent to launch evaluations, review traces, and write a new variant of itself.

He also showed a sanitized production view covering seven weeks of nightly CI jobs that benchmark candidate variants nightly.

Those jobs hover around 66 percent on the tracked task sets, and one break last night required Arya to fix itself before the talk.

Scoring is where he spends most of his time, with two modes: normative pass or fail and relativistic style comparisons.

The task library now holds 886 tasks categorized by levels and exposed to product teams for relevance checks.

Tasks range from single instructions to multi-turn personas that simulate users interacting with Arya in sequence.

The agent harness itself is deliberately simple and model agnostic, tested against CoreWeave inference and foundation model providers.

Context handling, compaction, and UI payload assembly are abstracted so YAML can spawn many parallel variants cheaply.

In the live run Arya converted a real production error, a missed weave.log SDK call in the sandbox, into a WBAF regression task and re-ran prod versus candidate.

He has not written code directly in about eight months, directing Claude to implement changes while he focuses on guardrails.

The limitation remains operational: simulation is expensive, evaluations take time to run, and the gap between offline scores and live behavior still needs manual review.

Outside Weights & Biases, others chase the same loop, with Nous Research describing its hermes-agent as a self-improving AI agent that grows with you.

For now Arya also trains models on H200s on CoreWeave infrastructure and runs auto-research on Karpathy's nano chat without leaving the Weave platform.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.