About Me

Hi, my name is Ruilin Yao (姚瑞霖). I am a PhD student at Wuhan University of Technology and a joint PhD student at the Institute of Automation, Chinese Academy of Sciences (CASIA). My research lies at the intersection of visual perception and multimodal learning, with a particular focus on visual grounding and its extension to GUI grounding. Within vision-language models, I investigate multimodal reasoning and language-guided perception, alongside core visual perception problems in object detection and segmentation. I am especially interested in equipping GUI agents with reliable long-horizon interaction capabilities and advancing general-purpose visual agentic reasoning, enabling agents to robustly perceive, ground, reason, and act in complex visual and human–computer interaction environments.

Research Keywords

  • Visual Perception
  • Visual Agentic Reasoning
  • Multimodal Reasoning
  • Long-Horizon GUI Agents

Educations

  • Wuhan University of Technology
    PhD, Computer Science and Technology (2024.09 - 2028.06), Advisor: Shengwu Xiong
  • Institute of Automation, Chinese Academy of Sciences (CASIA)
    Joint PhD Program (2024.09 - 2028.06), Advisor: Jinqiao Wang
  • Wuhan University of Technology
    MSc, Software Engineering (2021.09 - 2024.06)
  • Wuhan University of Technology
    BSc, Information and Computing Science (2017.09 - 2021.06)

Honors and Awards

  • Young Elite Scientists Sponsorship Program by CAST — Doctoral Student Special Plan, 2025–2027
  • Outstanding Graduate Award, Top 20%, Wuhan University of Technology, Jun. 2021
  • Excellence Scholarship, 10/260, Wuhan University of Technology, 2021–2022

Selected Publications

  • Background Blurring Matters: Improving Visual Grounding by Merging Text-Irrelevant Tokens Mr-Bigworth/ToB
    Ruilin Yao, Shengwu Xiong, Shanshan Yang, Tianyu Zou, Shili Xiong, and Yi Rong.
    ECCV 2026, First Author
  • AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding Mr-Bigworth/AutoFocus
    Ruilin Yao et al.
    arXiv:2605.02630, 2026, First Author
  • Lens: Multi-level evaluation of multimodal reasoning with large language models Lens4MLLMs/LENS
    Ruilin Yao, Bo Zhang, Jirui Huang, Xinwei Long, Yifang Zhang, Tianyu Zou, Yufei Wu, Shichao Su, Yifan Xu, Wenxi Zeng, Zhaoyu Yang, Guoyou Li, Shilan Zhang, Zichan Li, Yaxiong Chen, Shengwu Xiong, Peng Xu, Jiajun Zhang, Bowen Zhou, David Clifton, Luc Van Gool.
    ICLR 2026, First Author
  • Balancing conservatism and aggressiveness: Prototype-affinity hybrid network for few-shot segmentation tianyu-zou/PAHNet
    Tianyu Zou, Shengwu Xiong, Ruilin Yao, and Yi Rong.
    ICCV 2025, Third Author
  • Mars2 2025 challenge on multimodal reasoning: Datasets, methods, results, discussion, and outlook
    P. Xu, S. Xiong, J. Zhang et al.
    ICCV 2025 Workshop, Competition organizer & core author
  • MAP: Parameter-efficient tuning for referring expression comprehension via multi-modal adaptive positional encoding
    Ruilin Yao, Yi Rong, Tianyu Zou, Bo Zhang, Jian Li, Shengwu Xiong, and Shili Xiong.
    ACM MM 2025 (Oral), First Author
  • Visual grounding with multi-modal conditional adaptation Mr-Bigworth/MMCA
    Ruilin Yao, Shengwu Xiong, Yichen Zhao, and Yi Rong.
    ACM MM 2024 (Oral), First Author
  • CTOD: Cross-attentive task-alignment for one-stage object detection Mr-Bigworth/CTOD
    Ruilin Yao, Yi Rong, Qiangqiang Huang, and Shengwu Xiong.
    IEEE TCSVT 2024, First Author

Projects & Service

  • Research Intern, Ant Group (inclusionAI) UI-Venus
    Mar. 2026 – Present
    Enhancing GUI agents and grounding through trajectory synthesis with MLLMs in the UI-Venus Team; responsible for mid-training and post-training for grounding tasks, and for developing CAPTCHA-handling capabilities for GUI agents.
  • General Visual-Dialog Dataset and Multimodal Perception & Reasoning Evaluation
    Collaborated with Tsinghua University and CASIA on dataset collection, annotation, and construction; evaluated open- and closed-source models on perception (e.g., detection and OCR) and reasoning tasks, leading to an ICLR 2026 first-author paper.
  • ICCV MARS2 Workshop organization
    Proposal writing, workshop application, invited speakers, on-site hosting; assisted competition operation and EvalAI-based platform development.
    Homepage: https://mars2workshop.github.io/iccv2025

Competitions

  • Winner (1st Place, Online) & 3rd Place (Offline), ATEC 2021 Tech Elite Challenge (Ant Group)