SonicArt-Intelligence v3.0

September 8, 2024 · 1 min read
projects

“SonicArt-Intelligence” is an intelligent and immersive music education system designed for global music learners. Its aim is to establish a new generation of digital music learning ecosystem and create a brand-new learning environment for global music learners.

UE5 Scene Gamified Interaction

Junhao Wu
Authors
Junhao Wu (he/him)

I am an undergraduate student majoring in Computer Science at Dalian University of Technology. My research interests focus on Embodied AI and Vision-Language-Action (VLA) models, with the core goal of improving the robustness and zero-shot generalization of robots in complex environments.

My recent research focuses on enhancing model generalization. I developed a self-supervised prompt learning framework to bolster the visual robustness of VLA models (submitted to NeurIPS 2026). Additionally, I proposed MGTSM, a meta-learning framework for Open-Vocabulary Pedestrian Attribute Recognition (OVPAR) that effectively bridges the gap between seen and unseen categories (under review at IEEE TMM). I also held an Algorithm Internship at SDIC Intelligence, where I conducted research on multimodal deepfake detection and improved the system’s ability to detect cross-modal inconsistencies in synthetic media.

Currently, as a Research Intern at the College of AI, Tsinghua University, I am exploring the application of World Models in human sequence prediction and building an effective embodied data generation system to support downstream imitation learning and policy training. In parallel, as an Algorithm Intern at Apex Intelligence, I work on Auto Research agents and evaluation benchmarks, with a long-term interest in self-evolving models and ASI.

In the future, I hope to develop reproducible and scalable methods to continuously enhance the reliability and generalization of VLA models in real-world settings.