The Vision-Force-Action Model
Swedish robotics startup Eka demonstrated a robotic claw trained using a novel vision-force-action (VFA) model. Unlike traditional robotic training that relies on either visual data OR physical simulation, Eka's system learns from the simultaneous correlation of what it sees, what it feels, and what it does.
The approach mimics human learning: when you pick up an egg, you simultaneously see the egg, feel its weight and fragility, and adjust your grip. Eka's VFA model creates the same multi-modal feedback loop for robots.
Results and Applications
The VFA-trained robotic claw shows remarkable dexterity:
- Handles fragile objects (eggs, glass) with 99.2% success rate — up from 94% for vision-only
- Adapts to unfamiliar objects without retraining — generalizes from training data
- Works in cluttered environments where vision-only systems fail
- Learns new manipulation tasks 5x faster than traditional reinforcement learning
The multi-modal approach is what makes this interesting. Most robotic AI focuses on a single sense — usually vision. But real-world manipulation requires integrating multiple senses simultaneously. Eka's insight — that force feedback is as important as visual perception — seems obvious in hindsight, but it's the kind of insight that only comes from building real hardware, not just simulators.