Mingchao Sun (孙铭超)
Researcher · Alibaba Group (AMap)
pushing hard && past the limit -> embrace the dream
about
Profile
Mingchao Sun is a Researcher at Alibaba Group (AMap). He previously interned at Microsoft Research Asia (MSRA) and Alibaba AI Labs, and received his B.Eng. and M.Eng. degrees from Shandong University, where he was advised by Prof. Baoquan Chen.
His research at AMap focuses on 3D reconstruction and generation, world models, and embodied intelligence. He is best known for PointCNN (NeurIPS 2018), a foundational convolution operator for point clouds co-authored with Prof. Yangyan Li.
highlights
News
- 2026.07 Released Abot-World and ABot-3DWorld — a universal world model that explores any 3D space. live demo → news →
- 2026.07 Released WorldRoamBench, an open-world benchmark probing the long-horizon stability of interactive world models. Benchmark homepage → news →
- 2026.06 Released ABot-Earth 0.5, a generative 3D model of the Earth. live demo → news →
- 2026.06 SocialNav, a human-inspired foundation model for socially-aware navigation — selected as a CVPR 2026 Best Paper Candidate! project → news →
- 2026.05 🏆 Won 1st place at the AGIBOT World Model Challenge. news →
- 2026.03 From Orbit to Ground accepted to CVPR 2026 Findings — generative city-scale photogrammetry from extreme off-nadir satellite imagery. project → news →
- 2026.02 Two papers accepted to CVPR 2026 — PointCNN++ (convolution on native points) and SocialNav (socially-aware navigation).
- 2026.02 CLoD-GS accepted to ICLR 2026 — continuous level-of-detail rendering built on 3D Gaussian Splatting.
- 2025.08 "Yunjing" (云境), AMap's immersive AI product — sharing its core capabilities and AMap's recent advances in digital-twin technology and applications. news →
selected work
Publications
sorted by date (newest first)
-
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
An action-conditioned video world model enabling real-time, long-horizon closed-loop interaction, trained on AAA games, simulators, and web video.
-
ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space
A universal multimodal 3D world model that turns text, image, and video inputs into high-fidelity, explorable 3D worlds.
-
ABot-N1: Toward a General Visual-Language Navigation Foundation Model
A step toward a general Visual Language Navigation foundation model, decoupling cognition from control via a slow-fast architecture guided by dual visual-language signals.
-
WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
A benchmark probing long-horizon stability of interactive world models.
-
ABot-EARTH 0.5: Generative 3D Earth Model
A generative 3D framework that synthesizes vast, seamless 3D environments from ubiquitous, geospatially referenced satellite imagery.
-
POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation
The first closed-loop benchmark for real-world POI-goal navigation, paired with a Brain-Action Framework that couples reasoning with continuous waypoint prediction.
-
Abot-n0: Technical Report on the VLA Foundation Model for Versatile Embodied Navigation
A unified Vision-Language-Action (VLA) foundation model that achieves a "Grand Unification" across 5 core tasks.
-
From Orbit to Ground: Generative City Photogrammetry from Extreme Off-Nadir Satellite Images
Generative city-scale photogrammetry from extreme off-nadir satellite imagery.
-
PointCNN++: Performant Convolution on Native Points
A performant convolution operator that works directly on native point clouds.
-
SocialNav: Training a Human-Inspired Foundation Model for Socially-Aware Embodied Navigation
A human-inspired foundation model for socially-aware embodied navigation.
CVPR 2026 Best Paper Candidate -
CLoD-GS: Continuous Level-of-Detail via 3D Gaussian Splatting
A framework that integrates a continuous LoD mechanism directly into a 3DGS representation.
-
DO-Conv: Depthwise Over-Parameterized Convolutional Layer
A depthwise over-parameterized convolutional layer that boosts 2D CNN backbones.
-
DeepPipes: Learning 3D Pipelines Reconstruction from Point Clouds
Reconstructing 3D pipeline structures from point clouds.
-
Mutual Information Maximization in Graph Neural Networks
A mutual-information maximization objective for graph neural networks.
-
PointCNN: Convolution on X-Transformed Points
A foundational convolution operator for deep learning on point clouds.
2021 WAIC Youth Outstanding Paper Award
experience
Experience
- Alibaba Group — AMap, Visual Technology Center 2021.11 – present
Researcher
HD Map · 3DGS · World Models
- Alibaba Group · Alibaba Cloud, Video Cloud & DingTalk 2020.07 – 2021.11
Algorithm Engineer
Face & Video Algorithms
- Alibaba Group · AI Labs 2018.08 – 2020.07
Research Intern
Autonomous Driving & V2X
- Microsoft Research Asia (MSRA) 2016.08 – 2017.04
Research Intern
Machine Learning & BI
- Shandong University 2017.09 – 2020.07
M.Eng. · Computer Science & Technology · advised by Prof. Baoquan Chen
- Shandong University 2013.09 – 2017.07
B.Eng. · Software Engineering
openings
Join the team
I'm recruiting self-motivated research interns passionate about 3D vision, 3D Gaussian Splatting, and generative world models. Interested? Reach out with your CV via WeChat or email — happy to chat!