Junhao Wu 🔥

Junhao Wujunhao wu

(he/him)

I am an undergraduate student majoring in Computer Science at Dalian University of Technology. My research interests focus on Embodied AI and Vision-Language-Action (VLA) models, with the core goal of improving the robustness and zero-shot generalization of robots in complex environments.

My recent research focuses on enhancing model generalization. I developed a self-supervised prompt learning framework to bolster the visual robustness of VLA models (submitted to NeurIPS 2026). Additionally, I proposed MGTSM, a meta-learning framework for Open-Vocabulary Pedestrian Attribute Recognition (OVPAR) that effectively bridges the gap between seen and unseen categories (under review at IEEE TMM). I also held an Algorithm Internship at SDIC Intelligence, where I conducted research on multimodal deepfake detection and improved the system’s ability to detect cross-modal inconsistencies in synthetic media.

Currently, as a Research Intern at the College of AI, Tsinghua University, I am exploring the application of World Models in human sequence prediction and building an effective embodied data generation system to support downstream imitation learning and policy training. In parallel, as an Algorithm Intern at Apex Intelligence, I work on Auto Research agents and evaluation benchmarks, with a long-term interest in self-evolving models and ASI.

In the future, I hope to develop reproducible and scalable methods to continuously enhance the reliability and generalization of VLA models in real-world settings.

Download Resume
Languages
100%
Chinese Native
80%
English Fluent
60%
Deutsch Basic
Experience

Research Intern

College of AI, Tsinghua University

College of AI, Tsinghua University

Under the supervision of Assistant Professor Yongchao Chen, I work on world-model-driven robot learning and embodied intelligence. My research interests include VLA, world action models, humanoid robot motion retargeting, and multimodal synthetic data generation. I am also responsible for the deployment, maintenance, and experimental operation of the lab’s dual-arm Franka platform. This work aims to combine the temporal prediction capability of world models with the semantic understanding of foundation models, improving robots’ action generation, policy generalization, and execution robustness in complex real-world environments.

Algorithm Intern

Apex Intelligence

Apex Intelligence

During my internship at Apex Intelligence, I worked on Auto Research agents and evaluation benchmarks. My work included Harness optimization, CoT analysis model development, and internal testing of Apex Research. I also contributed to several research and engineering deliverables, including CoRL and EMNLP submission support for Auto Research, Apex Research PR and technical report writing, and the construction, contribution, and full-scale testing of ASI benchmark tasks. These evaluations covered multiple models, including Claude Opus 4.8, Claude Sonnet, and MiniMax.

Overall, my work focused on automated research, agent evaluation, and platform capability building, reflecting my long-term interest in advancing self-evolving models and the realization of ASI.

Research Intern

Wangxuan Institute of Computer Technology, Peking University

Wangxuan Institute of Computer Technology, Peking University

Serving as a Research Intern at OV3Lab, under the supervision of Professor Jiahuan Zhou. We systematically analyzed the performance degradation of OpenVLA under visual distribution shifts and attributed it to environmental perturbations in visual representations. Based on this insight, we developed a self-supervised training framework for the vision encoder, introducing a small number of learnable prompt tokens optimized with a redundancy-reduction objective. Without requiring target-domain data or language model updates, the method effectively enhances visual invariance. Experiments on LIBERO-plus and LIBERO-pro demonstrate consistent robustness improvements under lighting, layout, and background shifts, along with strong zero-shot generalization capability.

Algorithm Intern

SDIC Intelligence Xiamen Information Co., Ltd (AI Institute)

SDIC Intelligence Xiamen Information Co., Ltd

During my internship at SDIC Intelligence, I worked on open-vocabulary pedestrian attribute recognition and multimodal content understanding. For OVPAR, our team developed three open-vocabulary benchmarks—OV-MSP60K, OV-RAP2, and OV-CelebPAR—to evaluate generalization from seen to unseen attributes. We also proposed MGTSM, a meta-learning framework that simulates distribution shifts to improve adaptation to unseen classes, achieving more balanced performance across ZSL and GZSL settings.

I also contributed to a multimodal deepfake detection system built around a multi-branch architecture that combines semantic, frequency-domain, and spatial-domain cues, using consistency constraints and contrastive learning to improve feature fusion. Evaluated on more than 100,000 samples, the model achieved 0.97 AUC and over 93% accuracy.

Experience

Research Intern

Wangxuan Institute of Computer Technology, Peking University

Serving as a Research Intern at OV3Lab, under the supervision of Professor Jiahuan Zhou. We systematically analyzed the performance degradation of OpenVLA under visual distribution shifts and attributed it to environmental perturbations in visual representations. Based on this insight, we developed a self-supervised training framework for the vision encoder, introducing a small number of learnable prompt tokens optimized with a redundancy-reduction objective. Without requiring target-domain data or language model updates, the method effectively enhances visual invariance. Experiments on LIBERO-plus and LIBERO-pro demonstrate consistent robustness improvements under lighting, layout, and background shifts, along with strong zero-shot generalization capability.

WicT

Algorithm Intern

SDIC Intelligence Xiamen Information Co., Ltd (AI Institute)

Contributed to a multimodal deepfake detection system that fuses semantic, frequency-domain, and spatial-domain features for social platform content moderation and evidence extraction. Worked with 100K+ data samples across multiple modalities; incorporated CLIP-based semantic features, a frequency-branch module (Freq-Net), and a spatial-branch module (CameraPrint), and applied mutual-consistency constraints and contrastive learning to improve fusion. Achieved 0.97 AUC on the validation set and 93%+ accuracy, delivering a ~8-percentage-point improvement over a baseline model.

MeiyaPico

Education

CS Undergrad

Dalian University of Technology

  • During my studies, I have developed a solid and well-rounded profile across academics, student leadership, and innovation competitions:

    - Academic Performance: Major in Computer Science and Technology with a weighted GPA of 93.02 over the first five semesters, ranking 2/155 in my major. I am also pursuing a minor in International Organizations and Global Governance.

    - Honors & Awards: Recipient of the National Scholarship for two consecutive years; awarded Outstanding Student Model at Dalian University of Technology (top 10 university-wide).

    - Student Leadership: Serving as a member of the Student Union Presidium in the School of Computer Science, with strong experience in organization, coordination, and team management.

    - English Proficiency: CET-4: 618(Oral: Excellent), CET-6: 585 (Oral: Good).

    - Innovation & Competitions: Actively engaged in research and innovation competitions, earning 4 national-level awards and 40+ awards at various levels, with hands-on experience in project execution and delivering outcomes.

Projects

MindCare v6.0

MindCare v6.0

MindCare is an intelligent diagnosis system focused on the mental health of teenagers. It integrates virtual reality and cloud technology to establish a local-cloud collaborative architecture.

InnerVoice Guardian v2.0

InnerVoice Guardian v2.0

“InnerVoice Guardian” is an innovative artificial intelligence-based mental health support platform specifically designed to meet the increasing mental health needs of teenagers.

Events
Outstanding Student Leader featured image

Outstanding Student Leader

An honor exclusively awarded to 10 outstanding student leaders across School, recognizing excellent performance in student work and remarkable ability to serve students.

avatar
Junhao Wu
Marathon featured image

Marathon

I am a marathon enthusiast and have participated in multiple half and full marathon races. Recently finished the 2025 Guangzhou Marathon.

avatar
Junhao Wu
Outstanding Student Presentation Speech featured image

Outstanding Student Presentation Speech

Invited as an Outstanding Student Model(1/10) to deliver a university-wide live speech, sharing my journey of perseverance and serving as a role model to fellow students.

avatar
Junhao Wu
Awards
First Prize, China Undergraduate Mathematical Contest in Modeling (CUMCM)
China Undergraduate Mathematical Contest in Modeling ∙ 2025
Outstanding Student Model (Top 1/10)
Dalian University of Technology ∙ Oct 2025
Meritorious Winner (M Award), Mathematical Contest in Modeling (MCM/ICM)
COMAP, USA ∙ May 2025
Third Prize, Huawei ICT Competition
Huawei ICT, China ∙ Mar 2025
Second Prize, 21st Century Cup National English Speaking Competition
China Daily ∙ Mar 2024
National Third Prize for Three Consecutive Years, National English Competition for College Students (NECCS)
National English Competition for College Students ∙ 2024, 2025, 2026
National Scholarship (2024, 2025)
The Central People's Government of the People's Republic of China ∙ 2024, 2025
Outstanding Merit Student (Top 1/10)
Dalian University of Technology ∙ 2024, 2025
Skills & Hobbies
Skills
Leadership
Mathematical Modeling
Communication
Python / PyTorch
LaTeX
Hobbies
Marathon
Gym
Piano
Drum
Basketball