#Software Engineering

50 articles with this tag

Coding Agents Fail Rigorous Migration Tests
AI Research

Coding Agents Fail Rigorous Migration Tests

A new benchmark, SWE Refactor Bench, reveals that even frontier AI coding agents struggle to perform complete and correct whole-repository software migrations, highlighting a critical gap in current capabilities.

10 days ago
Warp's Safia Abdalla on Building Cloud AI Agents
Artificial Intelligence

Warp's Safia Abdalla on Building Cloud AI Agents

Safia Abdalla of Warp discusses building cloud AI agents, emphasizing developer workflow adaptation, complexity abstraction, and the power of an open API.

13 days ago
Rémi Louf: Agent Frameworks Are "Harmful"
Artificial Intelligence

Rémi Louf: Agent Frameworks Are "Harmful"

Rémi Louf, CEO of .txt, argues that current AI agent frameworks are "harmful" due to their reliance on manual intervention, advocating for event-driven systems and robust logging.

13 days ago
Software 3.0: The Next AI Paradigm Shift
AI Research

Software 3.0: The Next AI Paradigm Shift

A new paradigm, Software 3.0, is emerging, driven by context and reasoning, converging on databases, large models, and agents.

14 days ago
Andrej Karpathy Retires Vibe Coding for Agentic Engineering
AI Figures

Andrej Karpathy Retires Vibe Coding for Agentic Engineering

Andrej Karpathy coined the phrase 'vibe coding' in February 2025 to describe casual, low-scrutiny AI-assisted programming. By December 2025, 80 percent of his own code was AI-generated, and by April 2026 he had declared vibe coding passe at Sequoia Capital's AI Ascent event. What replaced it, a concept he calls agentic engineering, defines a harder discipline. Here is the 18-month arc of how his thinking changed.

20 days ago
Circleback CEO: Recording Meetings Will Be the Norm
Startup News

Circleback CEO: Recording Meetings Will Be the Norm

CircleBack CEO Ali Haghani discusses his AI note-taking platform, its use in candidate tracking, and the future of meeting recording in the age of AI.

24 days ago
Emulated Founders Detail Data Engine for Autonomous AI Engineers
Artificial Intelligence

Emulated Founders Detail Data Engine for Autonomous AI Engineers

Emulated co-founders Joseph Wang and Sid Patlollu break down why training truly autonomous software engineers requires multi-node real-cloud environments.

about 1 month ago
Vaibhav Gupta: Fighting Slop with Slop
Artificial Intelligence

Vaibhav Gupta: Fighting Slop with Slop

Vaibhav Gupta of Boundary discusses fighting software 'slop' by embracing it, detailing Boundary's unique approach to building a programming language without code reviews and at 'agent speed.'

about 1 month ago
AI Agents Revolutionize Scientific Software
Artificial Intelligence

AI Agents Revolutionize Scientific Software

AI agents are transforming scientific computing, speeding up software development and maintenance, but human oversight and long-term stewardship remain crucial.

about 1 month ago
Data Curve Launches DeepSWE Coding Benchmark
AI Research

Data Curve Launches DeepSWE Coding Benchmark

Data Curve introduces DeepSWE, a contamination-resistant coding benchmark designed to better evaluate AI coding agents on realistic, long-horizon software engineering tasks.

about 1 month ago
Arize CEO: AI Agents Will Automate Software Fixes
Artificial Intelligence

Arize CEO: AI Agents Will Automate Software Fixes

Arize CEO Jason Lopatecki discusses how AI agents are set to revolutionize software observability and debugging, enabling autonomous fixes and continuous self-improvement.

about 1 month ago
Alex Shaw: "Everything Is a Rollout" in AI Agent Evaluation
Artificial Intelligence

Alex Shaw: "Everything Is a Rollout" in AI Agent Evaluation

Alex Shaw from Lode Institute explains the Harbor framework, highlighting how agent development mirrors ML and requires empirical evaluation. Discover the tools and use cases for building and testing AI agents.

about 1 month ago
AI Agents as Supply Chain Actors: Patch Pilot's Security Model
Artificial Intelligence

AI Agents as Supply Chain Actors: Patch Pilot's Security Model

Moritz Johner of Form3 discusses the limitations of automated dependency patching tools and the security considerations of using AI agents with production code access, introducing their 'Patch Pilot' system.

about 2 months ago
Anthropic's Claude Code Creator on AI's Future
Artificial Intelligence

Anthropic's Claude Code Creator on AI's Future

Boris Cherny, Head of Claude Code at Anthropic, discusses the evolution of AI models, the future of software engineering, and Anthropic's focus on AI safety.

about 2 months ago
Bala Ramdoss on Generative UI for Agentic CX
Artificial Intelligence

Bala Ramdoss on Generative UI for Agentic CX

Bala Ramdoss of Amazon discusses generative UI, the critical layer between LLM output and product experience, emphasizing rendering contracts, streaming, and BFF patterns for agentic CX.

about 2 months ago
AI Agents Need Feature Flags for Safety, Says Engineer
Artificial Intelligence

AI Agents Need Feature Flags for Safety, Says Engineer

Backend engineer Sachin Gupta argues AI agents need specialized feature flags beyond traditional tools to manage their complex behaviors and mitigate risks.

about 2 months ago
Google DeepMind VP on AI's Role in Coding's Future
AI Research

Google DeepMind VP on AI's Role in Coding's Future

Google DeepMind's Benoit Schillings discusses how AI is reshaping software engineering, from code generation to scientific discovery.

about 2 months ago
AI's AI Reckoning Hits Software Engineers
Artificial Intelligence

AI's AI Reckoning Hits Software Engineers

Bloomberg Businessweek's Mark Milian discusses how AI is transforming software engineering, shifting roles from coding to AI direction and impacting the talent pipeline.

about 2 months ago
Addy Osmani: Own Your Verdict in the Age of AI Agents
Artificial Intelligence

Addy Osmani: Own Your Verdict in the Age of AI Agents

Addy Osmani, former Google Cloud AI Director, discusses the evolving role of engineers in the age of AI agents, emphasizing judgment, accountability, and ownership.

about 2 months ago
OpenAI Flags Major Flaws in SWE-Bench Pro
Artificial Intelligence

OpenAI Flags Major Flaws in SWE-Bench Pro

OpenAI's audit reveals approximately 30% of SWE-Bench Pro's coding tasks are flawed, prompting the company to retract its recommendation for the benchmark.

about 2 months ago
GitHub Fixes Repo Ownership
Technology

GitHub Fixes Repo Ownership

GitHub implemented a durable ownership system for all active repositories, archiving thousands of unmanaged ones and mandating ownership for new creations.

about 2 months ago
Octonous Streamlines AI Safety Work
Technology

Octonous Streamlines AI Safety Work

Mozilla.ai's AI Safety Engineer leverages Octonous to automate policy creation, monitor libraries, and aggregate research, boosting efficiency.

about 2 months ago
Duolingo's Angel Lee on AI Discernment vs. Approval
Artificial Intelligence

Duolingo's Angel Lee on AI Discernment vs. Approval

Angel Ortmann Lee from Duolingo discusses building AI systems for discernment, not approval, and the dangers of automation bias.

about 2 months ago
SWE-Marathon: Evaluating AI Coding Agents at Scale
AI Research

SWE-Marathon: Evaluating AI Coding Agents at Scale

Rishi Desai from Abundant AI introduces SWE-Marathon, a benchmark evaluating AI coding agents on billion-token scale tasks, revealing current limitations and the need for robust verification.

about 2 months ago
Wandero AI's Kalandadze on the 'Missing Layer' Post-Launch
Artificial Intelligence

Wandero AI's Kalandadze on the 'Missing Layer' Post-Launch

Wandero AI's CTO, Raphael Kalandadze, discusses the critical 'missing layer' of post-launch operations for AI agents, emphasizing the need for continuous monitoring and improvement loops.

2 months ago
Soheil Feizi on Continual Learning for AI Agents
AI Research

Soheil Feizi on Continual Learning for AI Agents

Soheil Feizi of RELAI explains the challenges and principles behind continual learning for AI agents, focusing on replayable, holistic, lifelong, and efficient improvements.

2 months ago
OpenAI's Bug Hunt: 18-Year-Old Flaw Found
Artificial Intelligence

OpenAI's Bug Hunt: 18-Year-Old Flaw Found

OpenAI uncovered two hidden bugs, including an 18-year-old software flaw, by analyzing crash data like an epidemiologist.

2 months ago
Cursor Brings AI Coding to iOS
Technology

Cursor Brings AI Coding to iOS

Cursor's new iOS app allows developers to manage AI coding agents from their phones, enabling development on the go.

2 months ago
Dominik Tornow: The Prompt is the Platform in AI
Artificial Intelligence

Dominik Tornow: The Prompt is the Platform in AI

Dominik Tornow of Resonate argues that AI development is shifting towards prompt engineering, making 'The Prompt is the Platform' a reality by 2026.

2 months ago
Microsoft Experts on Debugging Non-Deterministic AI Agents
Artificial Intelligence

Microsoft Experts on Debugging Non-Deterministic AI Agents

Microsoft experts Tisha Chawla and Susheem Koul discuss the challenges of debugging AI agents in production and introduce strategies for ensuring replayability and observability.

2 months ago
Angie Jones on Building Autonomous Engineering Orgs
Artificial Intelligence

Angie Jones on Building Autonomous Engineering Orgs

Angie Jones of Agentic AI Foundation discusses building autonomous engineering organizations, emphasizing AI as a collaborator and the importance of tailored integration.

2 months ago
Agents Building Agents: Nearform's AI Approach
Artificial Intelligence

Agents Building Agents: Nearform's AI Approach

Alfonso Graziano from Nearform explores how AI agents can build and improve other AI agents, detailing the 'Harness Engineering' methodology for reliable AI development.

2 months ago
OpenGov's Gabe De Mesa on Scaling AI Agents in Production
Artificial Intelligence

OpenGov's Gabe De Mesa on Scaling AI Agents in Production

Gabe De Mesa of OpenGov details how the company built and scaled its OG Assist AI agent, highlighting the use of Effect, A2A protocol, sandboxing, and developer velocity tools.

2 months ago
GitHub Copilot Harness Efficiency
Technology

GitHub Copilot Harness Efficiency

GitHub reveals its agentic harness matches model performance with superior token efficiency, supporting over 20 LLMs.

2 months ago
The Dawn of AI Agents: Building the First
Artificial Intelligence

The Dawn of AI Agents: Building the First

Experts discuss the evolution of AI agents, from early experimental tools to indispensable collaborators in software engineering.

2 months ago
Anthropic's Fiona Fung on Building AI-Pilled Engineering Teams
Artificial Intelligence

Anthropic's Fiona Fung on Building AI-Pilled Engineering Teams

Anthropic's Fiona Fung discusses how AI is transforming engineering teams, emphasizing initiative, growth mindset, and collaboration.

3 months ago
Nextdoor engineers build faster with Codex
Artificial Intelligence

Nextdoor engineers build faster with Codex

Nextdoor engineers are using OpenAI's Codex to accelerate development, enabling end-to-end feature building and faster debugging.

3 months ago
Anthropic Unleashes Claude Fable 5, Mythos 5
Artificial Intelligence

Anthropic Unleashes Claude Fable 5, Mythos 5

Anthropic launches Claude Fable 5 for general use and Mythos 5 for specialized cybersecurity, showcasing advanced capabilities with new safety measures and competitive pricing.

3 months ago
Code2LoRA: Repository Context without Overhead
AI Research

Code2LoRA: Repository Context without Overhead

Code2LoRA generates dynamic LoRA adapters for code LLMs, offering repository context without inference overhead and adapting to evolving codebases.

3 months ago
OpenClaw's Vincent Koc on 'Dark Factories' and AI Speed
Artificial Intelligence

OpenClaw's Vincent Koc on 'Dark Factories' and AI Speed

Vincent Koc of OpenClaw discusses the rapid acceleration of AI development, comparing it to the industrial revolution and highlighting OpenClaw's efficient "dark factory" approach.

3 months ago
Evaluating Coding Agents: Lessons from SWE-rebench
AI Research

Evaluating Coding Agents: Lessons from SWE-rebench

Ibragim Badertdinov from Nebius shares key lessons from evaluating coding agents using the SWE-rebench benchmark, highlighting the importance of real-world tasks, reliable verification, and cost-effectiveness.

3 months ago
Nvidia's Huang: AI Job Fears Are 'Nonsense'
Artificial Intelligence

Nvidia's Huang: AI Job Fears Are 'Nonsense'

Nvidia CEO Jensen Huang dismisses AI job loss fears as 'nonsense,' arguing AI actually drives demand for more software engineers.

3 months ago
Can LLMs Generate Enterprise-Quality Code?
Artificial Intelligence

Can LLMs Generate Enterprise-Quality Code?

Prasenjit Sarkar of Sonar discusses whether LLMs can generate enterprise-quality code, highlighting challenges and Sonar's AC/DC framework for agentic development.

3 months ago
Sakana AI: Finance Agents Take Shape
Technology

Sakana AI: Finance Agents Take Shape

Sakana AI is deploying AI agents to revolutionize financial operations, with engineers focusing on practical integration and enterprise-grade reliability.

3 months ago
Google DeepMind Explains AI Agent Building Struggles
AI Research

Google DeepMind Explains AI Agent Building Struggles

Philipp Schmid from Google DeepMind explains the core challenges senior engineers face when building AI agents, contrasting traditional engineering with agentic development.

3 months ago
Braintrust Cedes Coding to Codex
Artificial Intelligence

Braintrust Cedes Coding to Codex

Braintrust is dramatically speeding up its development cycle by integrating OpenAI's Codex, turning customer requests into code previews in minutes.

3 months ago
Cursor's RL Infrastructure for Training Composer
Artificial Intelligence

Cursor's RL Infrastructure for Training Composer

Cursor details its distributed infrastructure for training its AI coding model, Composer, using reinforcement learning on 'Fireworks'.

3 months ago
DeepMind's Scale: How Agents Run at Google
AI Research

DeepMind's Scale: How Agents Run at Google

Google DeepMind's KP Sawhney and Ian Ballantyne reveal how they run AI agents at scale, discussing the architecture, tools, and challenges involved in managing complex automated tasks.

3 months ago
LinkedIn Engineer Builds Community
tech

LinkedIn Engineer Builds Community

LinkedIn engineer Rishika builds community through mentorship and online content, extending her impact beyond her core role.

4 months ago
AI Agents Build Better AI
tech

AI Agents Build Better AI

LinkedIn Engineering details how AI agents are revolutionizing model development through automated, iterative refinement loops.

4 months ago