# Mango Rumors Exposed What Long Agents Forget _Erina Karati showed Project Paradox at AI Engineer and argued long-horizon agents need experimental loops that test rumor spread and memory provenance._ **Published:** 2026-09-28 **Source:** https://www.startuphub.ai/ai-news/artificial-intelligence/2026/mango-rumors-exposed-what-long-agents-forget --- Erina Karati, former engineer at Microsoft and [Supercell](/startups/supercell), brought a village that forgot to [AI Engineer](https://www.youtube.com/watch?v=x4e5O9zN0TE) and used it to argue that long-horizon agents need controlled experiments, not better prompts. A mango rumor broke it. Her team built [Project Paradox](https://www.youtube.com/watch?v=x4e5O9zN0TE) in [Supercell](/startups/supercell)'s AI Innovation Lab in Helsinki, where a 10-week sprint turned a Mafia-style social deduction game into a reusable multi-agent framework. The lab itself ran as a pilot with 16 people given 10 weeks to build AI games from scratch, and the public repository notes the build window as April 2025 to June 2025 by Karati and Arunachalam Manikandan. The architecture was intentionally stateful. Each agent carries a personality string that seeds every LLM call, plus an emotion vector of joy, sadness, fear, anger and disgust that shifts with events, a belief score for every other character that updates with each memory and decays when relationships go quiet, and a three-tier memory with seven items in RAM, a permanent cache for memories scoring 8 or above, and a per-character FAISS semantic store for everything else. Agents can move with intent, pick up and drop objects, react to events, and store conversations that change their beliefs and goals. For short runs it held. Ask Blossom to picnic and she plans the steps, grabs a pastry, walks to the spot, and answers in context. Over longer horizons the social consistency thinned. In Karati's demo an agent spreads a rumor about a mango sale, a second agent retells it, and after intervening events the player asks about mangoes and gets an answer that has lost source, hedged uncertainty, or the fact it was once a rumor stated as truth. The timeline that makes this talk possible matters. The closest public ancestor is the 2023 Stanford and Google Generative Agents paper where 25 agents lived in a Sims-like sandbox. Project Paradox refactored that idea for a game framework where inner state drives engine actions. Then, a few months before Karati's talk, Andrej Karpathy, former director of AI at Tesla and co-founder of [OpenAI](https://www.startuphub.ai/startups/openai), open-sourced AutoResearch, a compact Python framework released on GitHub in March 2026 that compresses an LLM training harness into about 630 lines. Karati frames that release as the missing loop. AutoResearch in Karpathy's own overnight run completed 126 experiments and reduced validation loss from 0.9979 to 0.9697 bits per byte without human intervention. Two days of unattended runs pushed about 700 changes and cut time-to-GPT-2 by 11 percent, and early adopters like Shopify CEO Tobi Lütke reported a 19 percent validation gain shortly after release. Karati is not proposing to drop that exact trainer into games. She proposes to borrow its shape and apply it as a meta system outside the village. The loop she described is tight and skeptical. Define a controlled scenario, run the simulation, collect structured traces of observations, conversations, memory writes, retrievals and belief updates, score against ground truth, propose a small policy edit, rerun, keep only what improves. The village keeps local perspectives, no common memory database, information travels only when agents communicate it. The auto research layer reads full traces and asks whether society-level behavior got better. Scenario design is the control surface. Let agents wander and you get charming demos you cannot measure. Fix a test and you can. One agent learns the bakery will close tomorrow. Do the right agents learn it and repland? One agent hears agent C might leave. Does might become is leaving when retold? A planned route blocks. Do agents update and tell each other? These suits replace vibes with repeatable runs. The scorecard has to stay balanced. Reach for diffusion, source retention for provenance, uncertainty preservation and false-sure rate for rumors, action consistency and time to replan for planning, containment for privacy. Optimize diffusion alone and agents overshare. Optimize recall alone and you get noisy or stale memories. A single vague agent quality metric hides the interesting failures, Karati warned, so the loop must ratchet forward only when the whole scorecard improves and guardrails hold. The other lesson is to keep the editable surface small. Freeze the harness, scenarios and metrics. Expose only memory writing policy, retrieval policy, communication prompts, belief and trust rules, source attribution and replanning triggers. If source attribution disappears, require source preserved in memory writes. If rumors harden into facts, store firsthand versus secondhand confidence and require hedging when retelling uncertain claims. If public facts stay local, classify public facts and make agents proactively share. The point is not rewriting the app, it is searching a controlled policy space. Karati is careful about what she does not claim. Without repeated controlled loops she would not say the system generally improved, only that this is the right surface to expose because it is small enough to control and rich enough to move social behavior. Rollback is not optional. A change that spreads public facts faster can also leak private information. More recall can mean more stale use. Memory alone does not fix it either. You need provenance, separation of raw episodic memories from current beliefs, and tests that force agents to act on what they actually know. The implication reaches beyond villages. Support agents need to know which policy update supersedes an older answer and where it came from. Assistants need to remember commitments and handle corrections. Research agents need citations and contradiction handling. Coding agents need context across issues and changing requirements. All maintain state that affects future action, and all fail the same way when state is just text in a vector store. If you take one recipe from the talk, it is freeze the harness, define scenarios, log traces, score behavior, and allow only small policy edits that survive measurement. Long horizons do not get reliable by watching one good demo. They get reliable when a system runs experiments on itself overnight and keeps receipts. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.