Who reported this story?

This story was reported by arXiv cs.RO.

Robotics

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes

Robos News Newsroom

Editorial Desk

2026-07-01 · 2 min read

Published July 1, 2026 · Category: Robotics

Overview

arXiv:2604.04834v2 Announce Type: replace-cross Abstract: Robotic Vision-Language-Action (VLA) models generalize well for open-ended manipulation, but their perception is fragile under sensing-stage degradations such as extreme low light, motion blur, and black clipping. We present E-VLA, an event-augmented VLA framework that improves manipulation robustness when conventional frame-based vision becomes unreliable. Instead of reconstructing images from events, E-VLA directly leverages motion and structural cues in event streams to preserve semantic perception and perception-action consistency under adverse conditions. We build an open-source teleoperation platform with a DAVIS346 event camera and collect a real-world synchronized RGB-event-action manipulation dataset across diverse tasks and illuminations. We also propose lightweight, pretrained-compatible event integration strategies and study event windowing for stable deployment. Experiments show that even a simple parameter-free fusion, i.e., overlaying accumulated event maps onto RGB images, could substantially improve robustness in dark and heavy-blur scenes: on Pick-Place at 20 lux, success increases from 0% (image-only) to 60% with overlay fusion and to 90% with our event adapter; under severe motion blur (1000 ms-exposure proxy), Pick-Place improves from 0% to 20-25%, and Sorting from 5% to 32.5%. Overall, E-VLA provides systematic evidence that event-driven perception can be effectively integrated into VLA models, pointing toward robust embodied intelligence beyond conventional frame-based imaging. Code and dataset will be available at https://github.com/JJayzee/E-VLA.

Source

Originally published at arxiv.org.

Source: https://arxiv.org/abs/2604.04834

Robos News Newsroom

Robos News covers markets, crypto and commodities for Asia & the Middle East — tier-1 desk research, AI-driven analysis, institutional-grade data. Tip our newsroom: [email protected]

Email the newsroom →

Disclaimer: This article is for informational purposes only and does not constitute investment advice. Data may be delayed up to 15 minutes. Past performance is not indicative of future results. Consult a licensed financial advisor before making investment decisions.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes

Overview

Source

Related Articles

Related Stories

Overview

Source

Related Articles

Related Stories

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

Multi-Robot Coordination for Planning under Context Uncertainty

Hierarchical 3D Scene Graph Construction and Belief-based Planning for Semantic Navigation

LDHP: Library-Driven Hierarchical Planning for Non-prehensile Dexterous Manipulation

Cookie Preferences