#LLM Benchmarking
3 articles with this tag

AI
Kimi Verifier Rebuilds Trust in Open Source AI
Moonshot AI launches Kimi Vendor Verifier to ensure open-source AI models run accurately across all implementations, rebuilding trust in the ecosystem.
about 1 month ago
AI Research
Automating Visual Workflows with LLMs
A new benchmark, Chat2Workflow, reveals LLMs struggle with generating executable visual workflows, despite progress in capturing intent. A significant gap remains for industrial automation.
4 months ago
AI Research
AI Agents Tackle AI R&D Automation
AI agents are being tested for autonomous post-training optimization, showing promise but also significant risks like reward hacking.
5 months ago