Dafeng Wei

Researcher · Finch, Embodied Research Center, AgiBot (智元机器人)

I work on Vision-Language Models (VLM) and Vision-Language-Action models (VLA) for embodied AI — teaching robots to reason about the world and act in it.

Previously, I was a senior algorithm engineer on the autonomous driving team at Li Auto, where I built multimodal large models for driving scenarios (on-vehicle VLM deployment related to DriveVLM, and multimodal retrieval systems for the data loop), and before that an algorithm engineer at ByteDance. I received my M.S. from Shanghai Jiao Tong University, advised by Prof. Hongtao Lu.

Portrait of Dafeng Wei

News

Publications

* denotes equal contribution. Also see my Google Scholar profile.

2025

GenieReasoner overview figure

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

Yi Liu, Sukai Wang, Dafeng Wei, Xiaowei Cai, Linqing Zhong, Jiange Yang, Guanghui Ren, Jinyu Zhang, Maoqing Yao, Chuankang Li, Xindong He, Liliang Chen, Jianlan Luo

arXiv preprint, 2025

GenieReasoner jointly optimizes embodied reasoning and action execution, with the ERIQ reasoning benchmark (6k+ QA pairs) and FACT, a flow-matching action tokenizer.

AgiBot World Colosseo overview figure

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

AgiBot-World-Contributors (incl. Dafeng Wei)

IROS 2025 🏆 Best Paper Award Finalist · IEEE T-RO 2026

An open platform with 1M+ real-robot trajectories across 217 tasks, and GO-1, a generalist policy built on latent action representations.

BEV-TSR overview figure

BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving

Tao Tang*, Dafeng Wei*, Zhengyu Jia, Tian Gao, Changwei Cai, Chengkai Hou, Peng Jia, Kun Zhan, Haiyang Sun, Jingchen Fan, Yixing Zhao, Fu Liu, Xiaodan Liang, Xianpeng Lang, Yang Wang

AAAI 2025

Retrieving complex driving scenes with free-form text queries directly in BEV feature space.

2022

Efficient SlowFast thumbnail

Efficient Dual Attention SlowFast Networks for Video Action Recognition

Dafeng Wei, Ye Tian, Liqing Wei, Hong Zhong, Siqian Chen, Shiliang Pu, Hongtao Lu

Computer Vision and Image Understanding (CVIU), 2022

Lightweight two-stream video networks with a cross-modality dual attention fusion module (CMDA).

2021

Active learning pipeline figure

Towards Dynamic and Scalable Active Learning with Neural Architecture Adaption for Object Detection

Fuhui Tang*, Dafeng Wei*, Chenhan Jiang, Hang Xu, Andi Zhang, Wei Zhang, Hongtao Lu, Chunjing Xu

BMVC 2021

Active learning for object detection that adapts network architecture as the labeled pool grows; deployed on large-scale autonomous-driving data at Huawei.

2020

TableCell dataset overview from the ICPR poster

Image-based Table Cell Detection: a Novel Table Structure Decomposition Method with New Dataset

Dafeng Wei, Hongtao Lu, Yi Zhou, Kai Chen

ICPR 2020

Detecting table cells as objects to recover table structure, with TableCell — an open dataset of 170K cell-level annotations.

Experience

Internships

Education

Miscellaneous