Tag: computervision
3 posts
Qwen-Drive 1.0: Turning a General VLM into a Driving Foundation Model
Qwen-Drive-1.0 turns Qwen3.5-4B into a unified autonomous-driving model for 3D perception, scene understanding, reaso...
LocateAnything Explained: Parallel Box Decoding and how it makes visual grounding faster and more precise
A review of LocateAnything, an NVIDIA vision-language model that treats each bounding box as one atomic unit and deco...
Gamma-World: Simplex Agent Encoding and Hub Attention for Multi-Agent World Models
A review of Gamma-World, NVIDIA's generative multi-agent world model that produces shared, action-controllable video ...