AI Engineer · AI
From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI
Feed a model one hour of video and roughly a million visual tokens go in, yet the loss lands on about two percent of them, since the only ground truth is a transcript or a few labeled frames. Armen Aghajanyan calls that a humongous waste, and predicting every pixel treats a background pixel with the same weight as a gripper tip or a contact point. Perceptron's answer is a perceptive objective that learns which percepts will matter, rather than ha

Introductie van de bron.
AI Engineer