I am a Full-Time Research Assistant at MMLab, The Chinese University of Hong Kong, advised by Prof. Xiangyu Yue, and an undergraduate student at Southern University of Science and Technology (SUSTech). My research focuses on multimodal large language models for fine-grained audiovisual understanding, temporal grounding, and physical reasoning.
Since August 2026, I have also led the Foundation Models team at Omni-Intelligence. We are a team of approximately four developing a 7B-scale Brain-Language Model (neural LLM); this work is ongoing.
This GitHub is where I share research code, experiments, and software I build for day-to-day use.
- Multimodal LLMs for video understanding, dense audiovisual captioning, and temporal grounding
- Audio-essential physical reasoning through audiovisual reward design
- Brain-Language Models that connect neural signals and language
- Data curation, supervised fine-tuning, reinforcement learning, and evaluation for foundation models
- AVTIME: Reinforcing Long-Video Understanding with Bidirectional Time-Semantic Consistency Reward — Under Review, 2026.
Audiovisual LLMs for dense video captioning and temporal grounding, with AVTime-50K providing approximately 50K videos and 1.87M timestamped events. - AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward — arXiv, 2026.
Fine-grained audiovisual captioning with 100K aligned video-caption pairs, detail-aware GRPO rewards, and atomic-fact evaluation. - Omni-Physics: Reinforcing Audio-Essential Physical Understanding via Audiovisual Reward — Under Review, 2026.
Studying audiovisual reward design for physical reasoning that requires audio evidence. - MindPilot: Closed-loop Visual Stimulation Optimization for Brain Modulation with EEG-guided Diffusion — ICLR 2026 (Poster).
More publications are listed on Google Scholar.
- Battuta
Native keyboard and mouse sound applications for macOS and Windows. The macOS app uses SwiftUI, preloaded PCM audio, and idle suspension, with custom sound packs and local typing analytics. - sustech-cli
A TypeScript CLI and local MCP server for SUSTech services, with structured JSON/JSONL output, tool discovery, guarded write workflows, and native credential storage. - aitimeline
Conference timeline browser built with Next.js for tracking AI venues and deadlines. - CS203Bproject_IntelligentScissors
Interactive image segmentation course project implemented in Java.
- Personal page: wormforce.net/members/mingyang-wu
- Google Scholar: Publications
- GitHub: @aprylewu
- Email: mingyangwu@cuhk.edu.hk