Hello, I'm

Zihua Wang 王梓骅

Ph.D. Candidate in Software Engineering

Southeast University (东南大学)

Efficient & High-Quality Vision-Language Generation

Zihua Wang

About Me

I am a Ph.D. candidate at the School of Computer Science and Engineering, Southeast University, advised by Prof. Yu Zhang. Previously, I obtained my bachelor's degree from the School of the Gifted Young, University of Science and Technology of China (USTC).

Currently, I am a research intern at Alibaba Tongyi Lab - Mobile Agent Team, working on GUI Grounding Agent.

My research focuses on efficient and high-quality vision-language generation, including:

  • Speculative decoding for multimodal/VLA inference acceleration
  • GUI grounding & agent with RL strategies
  • Image captioning & Visual question answering

Contact

Skills

Python PyTorch Skills Development Agent Development IELTS 7.0 CET-4/6

News

2026 2 papers accepted at ACM MM 2026 (CCF-A Conference)!
2026 Paper accepted at IEEE Transactions on Multimedia (CCF-A Journal)!
2026 Paper accepted at AAAI 2026 (CCF-A Conference)!
2025.12 Joined project: Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents at Alibaba Tongyi Lab.

Selected Publications

= First author   |   Full list on Google Scholar

TMM 2026CCF-A Journal, JCR Q1

Adaptively Clustering Neighbor Elements for Image-Text Generation

Zihua Wang, Xu Yang, Haiyang Xu, Hanwang Zhang, Ming Yan, Fei Huang, Yu Zhang

Proposes a unified adaptive neighbor-element clustering attention that groups related visual and textual units to improve coherence and fidelity in image-text generation.

AAAI 2026CCF-A Conference

Efficient and Effective In-context Demonstration Selection with Coreset

Zihua Wang, Jiarui Wang, Haiyang Xu, Ming Yan, Fei Huang, Xu Yang, Xiu-Shen Wei, Siya Mi, Yu Zhang

Introduces a coreset-based dual retrieval algorithm that efficiently selects compact and informative in-context demonstrations for few-shot learning.

ACM MM 2026CCF-A Conference

FLASH: Latent-Aware Semi-Autoregressive Speculative Decoding for Multimodal Tasks

Zihua Wang, Ruibo Li, Haozhe Du, Joey Tianyi Zhou, Yu Zhang, Xu Yang

A latent-aware semi-autoregressive speculative decoding framework that compresses redundant visual tokens and generates drafts in parallel, accelerating multimodal generation.

ACM MM 2026CCF-A Conference

Open-Loop Planning, Closed-Loop Verification: Speculative Verification for VLA

Zihua Wang, Zhitao Lin, Ruibo Li, Yu Zhang, Xu Yang, Siya Mi, Xiu-Shen Wei

Proposes a speculative verification framework using a heavy VLA as a low-frequency macro-planner and a lightweight verifier for high-frequency closed-loop checking in embodied control.

TPAMI 2026CCF-A Journal · Under Review

Efficient Multimodal In-Context Learning via Coreset Retrieval and Speculative Decoding

Zihua Wang, Xu Yang, Xiu-Shen Wei, Siya Mi, Yu Zhang, Xin Geng

Introduces an efficient demonstration selection strategy balancing diversity and relevance via cluster-pruning coreset, plus a tailored speculative decoding scheme for acceleration.

Pattern RecognitionCCF-B Journal · Under Review

Exploring Patterns and Semantics in Non-Autoregressive Image Captioning

Zihua Wang, Xu Yang, Haiyang Xu, Ming Yan, Ji Zhang, Fei Huang, Yu Zhang

Deconstructs non-autoregressive generation into patterns and semantics, proposing label selection and image pre-fusion with joint training to close the quality gap with autoregressive models.

Neurocomputing 2025CCF-C Journal, JCR Q1

DVC2: Deep Video Cascade Clustering from Video Structures

Zihua Wang, Siya Mi, Yu Zhang

A hierarchical deep clustering framework leveraging spatial-temporal dependencies for interpretable and temporally stable video structure discovery.

ProjectAlibaba Tongyi

Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents

Haiyang Xu, ... Zihua Wang, ... Ming Yan

Multi-platform GUI agent framework for autonomous mobile device interaction.

PRCV 2024

Efficient Multi-modal Human-centric Contrastive Pre-training with a Pseudo Body-structured Prior

Yihang Meng, Hao Cheng, Zihua Wang, Hongyuan Zhu, Xiuxian Lao, Yu Zhang

ACL 2023

Transforming Visual Scene Graphs to Image Captions

Xu Yang, Jiawei Peng, Zihua Wang, Haiyang Xu, Qinghao Ye, Chenliang Li, et al.

FCS 2023

Person Video Alignment with Human Pose Registration

Yu Zhang, Zihua Wang, Jingyang Zhou, Siya Mi

Intern Experience

Alibaba Tongyi Lab — Mobile Agent Team

Dec 2025 — Present

Research Intern · GUI Agent & GUI Grounding

Working on multi-agent system, GUI grounding techniques, and RL-based strategies for GUI agents.

Alibaba DAMO Academy — NLP Group

May 2022 — Jul 2024

Research Intern · Multimodal Image-Text Models

  • Researched adaptive clustering attention methods for image-text semantic understanding
  • Developed coreset-based multi-modal in-context demonstration selection
  • Explored non-autoregressive visual text generation with joint training strategies

Education

Southeast University (东南大学)

Ph.D. in Software Engineering

School of Computer Science and Engineering

Mar 2023 — Jun 2027 (Expected)

Southeast University (东南大学)

M.S. in Electronic Information (Integrated M.S./Ph.D.)

School of Cyber Science and Engineering

Sep 2020 — Mar 2023

USTC (中国科学技术大学)

B.E. in Electronic Information Engineering

School of the Gifted Young (少年班学院)

Sep 2016 — Jun 2020

Honors & Awards

Outstanding Graduate Student
(三好研究生)
Huawei Scholarship
(华为奖学金)
Tang Zhongying Scholarship
(唐仲英奖学金)
SEU Second-Class Scholarship
(东南大学二等奖学金)
Silver Award, SEU Anniversary Entrepreneurship Competition
(校庆杯创新创业大赛银奖)
USTC Four-Star Volunteer
(中国科大四星级志愿者)

Services

Reviewer

ICLR, ACL, EMNLP, ACM MM, IEEE TMM, IEEE TCSVT, ACM TOMM, ACM MM Asia, FCS

First Aid Runner

2024, 2025 Nanjing Lishui Half Marathon

Projects

Interactive demos of my personal development projects. Click to explore.

歪脖 coding 中