#Computer Vision

50 articles with this tag

Unified Framework for Decision-Informed Future Prediction
AI Research

Unified Framework for Decision-Informed Future Prediction

DA-WAM unifies predictive representation learning and action-conditioned future modeling for safer autonomous driving, outperforming existing methods on key benchmarks.

9 days ago
Claude's Corner: AxionOrbital Space - Seeing Through the Clouds
Claude's Corner

Claude's Corner: AxionOrbital Space - Seeing Through the Clouds

AxionOrbital Space translates raw SAR satellite data into photorealistic optical imagery in 0.06 seconds, making Earth observation useful regardless of clouds or darkness. Their one-step diffusion model outperforms every published benchmark on MSAW, and they occupy a competitive category with no other commercial entrant.

10 days ago
Reactor's Ahmed Ahres on Real-Time Interactive Video
Artificial Intelligence

Reactor's Ahmed Ahres on Real-Time Interactive Video

Ahmed Ahres of Reactor discusses the transformative potential of real-time interactive video, moving beyond static generative models to dynamic, programmable content.

11 days ago
Krea.ai Details Krea 2 Image Model Training
AI Research

Krea.ai Details Krea 2 Image Model Training

Sangwha Lee of Krea.ai details the rigorous data curation and training process behind the Krea 2 image generation model, emphasizing stylistic diversity and efficiency.

11 days ago
Runway Research Unveils Autoregressive-to-Diffusion VLMs for Faster Visual AI
AI Video

Runway Research Unveils Autoregressive-to-Diffusion VLMs for Faster Visual AI

12 days ago
Runway General World Models Advance AI Simulation
AI Video

Runway General World Models Advance AI Simulation

Runway Research has launched a new research initiative called General World Models (GWM), aiming to create AI systems that can understand and simulate real-world visual dynamics and interactions.

12 days ago
Marionette: Decoupling Appearance from State
AI Research

Marionette: Decoupling Appearance from State

The Marionette world model decouples visual synthesis from geometric state, enabling controllable, long-horizon interactive simulations with explicit state repair.

12 days ago
MMDiff: Auditing and Steering MLLMs
AI Research

MMDiff: Auditing and Steering MLLMs

MMDiff, a new multimodal model-diffing framework, enables granular control and understanding of MLLMs by isolating and manipulating specific behavioral features.

18 days ago
Nvidia's Chris Alexiuk on Compression at the Edge
Semiconductors

Nvidia's Chris Alexiuk on Compression at the Edge

Nvidia's Chris Alexiuk discusses the critical role of data compression in enabling powerful AI processing on resource-constrained edge devices.

23 days ago
Uber Eats Uses AI Agents to Enhance Food Photos at Scale
Artificial Intelligence

Uber Eats Uses AI Agents to Enhance Food Photos at Scale

Uber Eats' computer vision team details their AI agent system for enhancing food photos, focusing on closed-loop feedback and continuous learning.

about 1 month ago
TwelveLabs Builds Video Memory Layer
Artificial Intelligence

TwelveLabs Builds Video Memory Layer

James Le of TwelveLabs discusses the critical need for memory in video AI, detailing the limitations of current systems and introducing their 'Jockey' platform as a solution.

about 1 month ago
ABot-World-0: Real-time Video World Models
AI Research

ABot-World-0: Real-time Video World Models

ABot-World-0 introduces a real-time video world model for long-horizon agent interaction, achieving 16 FPS at 720P with an optimized inference stack.

about 1 month ago
DiTs Unlock Precise Regional Image Control
AI Research

DiTs Unlock Precise Regional Image Control

New 'appearance pointers' enable Diffusion Transformers to precisely control image generation regionally, matching SOTA performance without retraining.

about 1 month ago
AI Vision Test Reveals Model Hallucinations
Artificial Intelligence

AI Vision Test Reveals Model Hallucinations

A new benchmark, PerceptionBench, reveals that even advanced multimodal AI models struggle with basic visual perception, often guessing answers instead of truly seeing.

about 1 month ago
WCog-VLA: Bridging Foresight for Proactive Autonomy
AI Research

WCog-VLA: Bridging Foresight for Proactive Autonomy

WCog-VLA pioneers a dual-level framework, unifying semantic forecasting with generative world evolution to enable proactive autonomous driving, achieving SOTA on NAVSIM.

about 2 months ago
State of the Union: Local AI Focus
Artificial Intelligence

State of the Union: Local AI Focus

Nvidia, Osmantic, Roboflow, and EXO Labs discuss the critical shift towards local AI solutions and its implications.

about 2 months ago
CARLA-GS: Unified Corner-Case Synthesis
AI Research

CARLA-GS: Unified Corner-Case Synthesis

CARLA-GS offers a unified, modular pipeline for synthesizing photorealistic and physically consistent corner cases in autonomous driving simulation.

about 2 months ago
ELSA3D: Structured Text-3D Reasoning
AI Research

ELSA3D: Structured Text-3D Reasoning

ELSA3D revolutionizes unified 3D models with elastic semantic anchoring, achieving SOTA performance in 3D generation and captioning while halving computational costs.

about 2 months ago
AI Learns to Smell: The Science of Olfactory AI
Artificial Intelligence

AI Learns to Smell: The Science of Olfactory AI

AI is learning to smell thanks to companies like Osmo, which are building models to understand, predict, and design scents by mapping molecular structures to olfactory perception.

about 2 months ago
CamVLA: Unshackling Robot Control from Camera Calibration
AI Research

CamVLA: Unshackling Robot Control from Camera Calibration

CamVLA revolutionizes robot control by enabling policies to infer camera geometry, achieving robust, calibration-free manipulation from single RGB images.

about 2 months ago
VisionAId: On-Device Vision for the Visually Impaired
AI Research

VisionAId: On-Device Vision for the Visually Impaired

VisionAId transforms smartphones into real-time visual assistants for the visually impaired, leveraging on-device AI and few-shot learning for personalized object recognition and multimodal guidance.

about 2 months ago
Netflix Tackles AI Video Editing Challenges
Artificial Intelligence

Netflix Tackles AI Video Editing Challenges

Netflix is developing advanced AI tools, Vera and VOID, to enhance video editing precision and realism for creators.

2 months ago
How Adult Websites Use AI: Inside Pornhub's Tech Stack and the Porn AI Boom
Insights

How Adult Websites Use AI: Inside Pornhub's Tech Stack and the Porn AI Boom

How do adult websites use AI? A technical deep dive into the Pornhub AI stack: computer vision tagging, recommendation pipelines, content moderation, and neural upscaling that power modern adult tech.

2 months ago
OneCanvas: Unified 3D Scene Representation
AI Research

OneCanvas: Unified 3D Scene Representation

OneCanvas revolutionizes 3D scene understanding in VLMs by projecting multi-view features onto a unified equirectangular canvas, enabling efficient situated reasoning and SOTA performance.

2 months ago
Phase Dominance in AI Image Recognition
AI Research

Phase Dominance in AI Image Recognition

AI image classifiers exhibit a striking phase dominance for identity encoding, mirroring human vision principles, with architectural differences shaping its expression.

2 months ago
ActiveSAM: Efficient Open-Vocabulary Segmentation
AI Research

ActiveSAM: Efficient Open-Vocabulary Segmentation

ActiveSAM revolutionizes open-vocabulary semantic segmentation with a training-free framework that dynamically identifies relevant classes, boosting speed and accuracy while enhancing robustness for real-world AI.

2 months ago
HYDRA-X: Unifying Image & Video Tokenization
AI Research

HYDRA-X: Unifying Image & Video Tokenization

HYDRA-X, a novel Vision Transformer-based UMM, unifies image and video tokenization, enhancing editing consistency and performance through causal attention and latent-level manipulation.

3 months ago
Humanoids Learn Self-Other Distinction
AI Research

Humanoids Learn Self-Other Distinction

Humanoid robots now learn self-other distinction and build predictive self-models from sensory data, enabling better collaboration and task performance in human-robot environments.

3 months ago
Rethinking VLM Token Reduction
AI Research

Rethinking VLM Token Reduction

Reroute transforms VLM token reduction from irreversible pruning to recoverable routing, improving grounding performance without sacrificing efficiency.

3 months ago
Images as the New Reasoning Medium
AI Research

Images as the New Reasoning Medium

This paper introduces optical reasoning, enabling images to serve as the primary medium for LLM and MLLM reasoning, achieving higher token efficiency and competitive performance.

3 months ago
Beyond Observable Data: Imaginative Perception for VLMs
AI Research

Beyond Observable Data: Imaginative Perception for VLMs

Researchers introduce Imaginative Perception Tokens (IPTs) to enable VLMs to reason about unobserved spatial configurations, outperforming textual chain-of-thought.

3 months ago
Fei-Fei Li Clarifies 'World Models'
Investors News

Fei-Fei Li Clarifies 'World Models'

Fei-Fei Li offers a framework to define AI 'world models', distinguishing them from language models and tracing their roots to agent-environment interaction.

3 months ago
AdaCodec: Efficient Video MLLM Encoding
AI Research

AdaCodec: Efficient Video MLLM Encoding

AdaCodec revolutionizes video MLLMs by using predictive visual coding to drastically cut tokenization costs and latency, achieving superior performance at a fraction of the budget.

3 months ago
GPIC: Fueling Next-Gen Generative Models
AI Research

GPIC: Fueling Next-Gen Generative Models

The GPIC dataset, a 28 trillion pixel permissive image corpus, democratizes large-scale visual generative model research and commercialization.

3 months ago
LocateAnything: Parallel Decoding for Vision
AI Research

LocateAnything: Parallel Decoding for Vision

LocateAnything revolutionizes vision-language models with Parallel Box Decoding, boosting speed and accuracy in visual grounding and detection.

3 months ago
Uber Fights Bounding Box Errors
tech

Uber Fights Bounding Box Errors

Uber Engineering uses machine learning to automatically detect and correct bounding box annotation errors in video data, boosting ML model training quality.

3 months ago
AI Image Generation Reimagined: Channel-Wise Quantization
AI Research

AI Image Generation Reimagined: Channel-Wise Quantization

Channel-wise Vector Quantization (CVQ) redefines image tokenization, enabling autoregressive models like CAR to generate richer, more detailed images with state-of-the-art performance.

3 months ago
SpatioRoute VLM: Dynamic Prompting for Video QA
AI Research

SpatioRoute VLM: Dynamic Prompting for Video QA

SpatioRoute VLM revolutionizes zero-shot spatial video question answering with dynamic prompt routing, achieving SOTA without fine-tuning or 3D sensors.

3 months ago
TaskGround: Bridging Scene Context and Action
AI Research

TaskGround: Bridging Scene Context and Action

TaskGround revolutionizes household AI by enabling compact models to interpret complex scenes, infer task structures, and act effectively, drastically improving performance and reducing costs.

3 months ago
AI Learns to See, Hear, and Understand
Technology

AI Learns to See, Hear, and Understand

Multimodal AI analytics is enabling businesses to decode video, audio, and images, unlocking deeper insights from previously unstructured data.

3 months ago
LMPath: Semantics Supercharge UAV Search
AI Research

LMPath: Semantics Supercharge UAV Search

LMPath integrates language and vision models to create semantically-aware exploration priors for UAVs, dramatically improving search mission efficiency over traditional geometric methods.

4 months ago
Beyond RGB: Grounding Vision-Language on Raw Sensor Data
AI Research

Beyond RGB: Grounding Vision-Language on Raw Sensor Data

PRISM-VL advances vision-language models by grounding them in raw camera measurements, not just RGB, significantly improving performance on challenging visual tasks.

4 months ago
Claude's Corner: GrazeMate, Three Clicks to Move a Thousand Cows
Claude's Corner

Claude's Corner: GrazeMate, Three Clicks to Move a Thousand Cows

GrazeMate builds fully autonomous drone software that herds cattle across million-acre stations with three phone taps, using proprietary reinforcement learning trained on expert stockmanship to read and respond to real-time animal behavior. Founded by a 19-year-old Australian farmer, the company has $1.2M raised, 1.7 million acres under contract, and is expanding into California and Texas.

4 months ago
Codex AI Automates Complex Computer Tasks
Artificial Intelligence

Codex AI Automates Complex Computer Tasks

Codex AI demonstrates advanced capabilities, automating complex tasks across applications by interacting with computer interfaces.

4 months ago
Claude's Corner: Librar Labs, The AI Librarian That's Really a Data-Catalog Trojan Horse
Claude's Corner

Claude's Corner: Librar Labs, The AI Librarian That's Really a Data-Catalog Trojan Horse

Librar Labs looks like another YC W2026 SaaS, AI-powered school library management, until you look at the team and the technical claim under the hood. OpenAI / Scale / Palantir alums plus quantum physicists plus a 'self-healing database for unstructured data' don't build a school librarian assistant unless the school librarian is the wedge.

4 months ago
MLX Genmedia: Prince Canuma on On-Device AI
Artificial Intelligence

MLX Genmedia: Prince Canuma on On-Device AI

Prince Canuma of MLX Genmedia discusses the power of on-device AI, showcasing how MLX enables efficient deployment of AI models on Apple Silicon devices for vision and audio tasks.

4 months ago
Transformers Conquer Computer Vision
Artificial Intelligence

Transformers Conquer Computer Vision

Isaac Robinson from Roboflow explains how Transformers, once confined to NLP, have revolutionized computer vision, surpassing CNNs through massive pre-training and architectural innovation.

4 months ago
Black Forest Labs: FLUX and the Future of Visual AI
AI Research

Black Forest Labs: FLUX and the Future of Visual AI

Stephen Batifol of Black Forest Labs discusses FLUX, the company's visual AI model, and the future of generative AI with a focus on real-time generation and world models.

4 months ago
PhyCo: Bridging Physics and Video Generation
AI Research

PhyCo: Bridging Physics and Video Generation

PhyCo introduces a framework for physically consistent and controllable video generation, overcoming limitations of current diffusion models through physics-supervised fine-tuning and VLM-guided rewards.

4 months ago
X-WAM: Bridging Action and 4D Synthesis
AI Research

X-WAM: Bridging Action and 4D Synthesis

The X-WAM unified 4D world model revolutionizes robotics by integrating real-time action with high-fidelity 4D synthesis, achieving state-of-the-art benchmarks.

4 months ago