


default search action
Yuhang Zang
Person information
Refine list

refinements active!
zoomed in on ?? of ?? records
view refined list in
export refined list as
2020 – today
- 2025
[j2]Yuhang Zang
, Wei Li, Jun Han, Kaiyang Zhou, Chen Change Loy:
Contextual Object Detection with Multimodal Large Language Models. Int. J. Comput. Vis. 133(2): 825-843 (2025)
[c28]Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Ziyu Liu, Shengyuan Ding, Shenxi Wu, Yubo Ma, Haodong Duan, Wenwei Zhang, Kai Chen, Dahua Lin, Jiaqi Wang:
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model. ACL (Findings) 2025: 6547-6563
[c27]Yubo Ma, Jinsong Li, Yuhang Zang, Xiaobao Wu, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Haodong Duan, Jiaqi Wang, Yixin Cao, Aixin Sun:
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings. ACL (Findings) 2025: 19568-19580
[c26]Jiazi Bu, Pengyang Ling, Pan Zhang, Tong Wu, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang:
ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way. CVPR 2025: 12999-13008
[c25]Long Xing, Qidong Huang, Xiaoyi Dong, Jiajie Lu, Pan Zhang, Yuhang Zang, Yuhang Cao, Conghui He, Jiaqi Wang
, Feng Wu, Dahua Lin:
Conical Visual Concentration for Efficient Large Vision-Language Models. CVPR 2025: 14593-14603
[c24]Zihao Huang, Shoukang Hu, Guangcong Wang, Tianqi Liu, Yuhang Zang, Zhiguo Cao, Wei Li, Ziwei Liu:
WildAvatar: Learning In-the-wild 3D Avatars from the Web. CVPR 2025: 15963-15975
[c23]Junbo Niu, Yifei Li, Ziyang Miao, Chunjiang Ge, Yuanhang Zhou, Qihao He, Xiaoyi Dong, Haodong Duan, Shuangrui Ding, Rui Qian, Pan Zhang, Yuhang Zang, Yuhang Cao, Conghui He, Jiaqi Wang:
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? CVPR 2025: 18902-18913
[c22]Rui Qian, Shuangrui Ding, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang
:
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction. CVPR 2025: 24045-24055
[c21]Pengyang Ling, Jiazi Bu, Pan Zhang, Xiaoyi Dong, Yuhang Zang, Tong Wu, Huaian Chen, Jiaqi Wang, Yi Jin:
MotionClone: Training-Free Motion Cloning for Controllable Video Generation. ICLR 2025
[c20]Ziyu Liu, Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Haodong Duan, Conghui He, Yuanjun Xiong, Dahua Lin, Jiaqi Wang:
MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models. ICLR 2025
[c19]Zihan Liu, Shuangrui Ding, Zhixiong Zhang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang:
SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation. ICML 2025
[c18]Xilin Wei, Xiaoran Liu, Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Jian Tong, Haodong Duan, Qipeng Guo, Jiaqi Wang, Xipeng Qiu, Dahua Lin:
VideoRoPE: What Makes for Good Video Rotary Position Embedding? ICML 2025
[i74]Rui Qian, Shuangrui Ding, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang
:
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction. CoRR abs/2501.03218 (2025)
[i73]Beichen Zhang, Yuhong Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Haodong Duan, Yuhang Cao, Dahua Lin, Jiaqi Wang:
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning. CoRR abs/2501.03226 (2025)
[i72]Yifei Li, Junbo Niu, Ziyang Miao, Chunjiang Ge, Yuanhang Zhou, Qihao He, Xiaoyi Dong, Haodong Duan, Shuangrui Ding, Rui Qian, Pan Zhang, Yuhang Zang, Yuhang Cao, Conghui He, Jiaqi Wang
:
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? CoRR abs/2501.05510 (2025)
[i71]Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Ziyu Liu, Shengyuan Ding, Shenxi Wu, Yubo Ma, Haodong Duan, Wenwei Zhang, Kai Chen, Dahua Lin, Jiaqi Wang
:
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model. CoRR abs/2501.12368 (2025)
[i70]Xilin Wei, Xiaoran Liu, Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Jian Tong, Haodong Duan, Qipeng Guo, Jiaqi Wang
, Xipeng Qiu, Dahua Lin:
VideoRoPE: What Makes for Good Video Rotary Position Embedding? CoRR abs/2502.05173 (2025)
[i69]Yujie Zhou, Jiazi Bu, Pengyang Ling, Pan Zhang, Tong Wu, Qidong Huang, Jinsong Li, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Anyi Rao, Jiaqi Wang
, Li Niu:
Light-A-Video: Training-free Video Relighting via Progressive Light Fusion. CoRR abs/2502.08590 (2025)
[i68]Zihan Liu, Shuangrui Ding, Zhixiong Zhang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang
:
SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation. CoRR abs/2502.13128 (2025)
[i67]Ziyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang
:
Visual-RFT: Visual Reinforcement Fine-Tuning. CoRR abs/2503.01785 (2025)
[i66]Yibin Wang, Yuhang Zang, Hao Li, Cheng Jin, Jiaqi Wang:
Unified Reward Model for Multimodal Understanding and Generation. CoRR abs/2503.05236 (2025)
[i65]Jiazi Bu, Pengyang Ling, Yujie Zhou, Pan Zhang, Tong Wu, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang
:
HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance. CoRR abs/2504.06232 (2025)
[i64]Shengyuan Ding, Shenxi Wu, Xiangyu Zhao, Yuhang Zang, Haodong Duan, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Dahua Lin, Jiaqi Wang
:
MM-IFEngine: Towards Multimodal Instruction Following. CoRR abs/2504.07957 (2025)
[i63]Yibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang, Qinglin Lu, Cheng Jin, Jiaqi Wang:
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning. CoRR abs/2505.03318 (2025)
[i62]Ziyu Liu, Yuhang Zang, Yushan Zou, Zijian Liang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang:
Visual Agentic Reinforcement Fine-Tuning. CoRR abs/2505.14246 (2025)
[i61]Jiaer Xia, Yuhang Zang, Peng Gao, Yixuan Li, Kaiyang Zhou:
Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning. CoRR abs/2505.14677 (2025)
[i60]Yubo Ma, Jinsong Li, Yuhang Zang, Xiaobao Wu, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Haodong Duan, Jiaqi Wang
, Yixin Cao, Aixin Sun:
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings. CoRR abs/2506.04997 (2025)
[i59]Long Xing, Qidong Huang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Jinsong Li, Shuangrui Ding, Weiming Zhang, Nenghai Yu, Jiaqi Wang, Feng Wu, Dahua Lin:
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing. CoRR abs/2506.19848 (2025)
[i58]Jiaer Xia, Bingkui Tong, Yuhang Zang, Rui Shao, Kaiyang Zhou:
Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation. CoRR abs/2507.02859 (2025)
[i57]Zhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He, Jianfan Lin, Junsong Tang, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang:
SeC: Advancing Complex Video Object Segmentation via Progressive Concept Construction. CoRR abs/2507.15852 (2025)
[i56]Jinsong Li, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Jiaqi Wang, Dahua Lin:
Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models. CoRR abs/2508.00819 (2025)
[i55]Zeyi Sun, Ziyu Liu, Yuhang Zang, Yuhang Cao, Xiaoyi Dong, Tong Wu, Dahua Lin, Jiaqi Wang:
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience. CoRR abs/2508.04700 (2025)
[i54]Jiazi Bu, Pengyang Ling, Yujie Zhou, Yibin Wang, Yuhang Zang, Tong Wu, Dahua Lin, Jiaqi Wang:
DiCache: Let Diffusion Model Determine Its Own Cache. CoRR abs/2508.17356 (2025)
[i53]Zeyi Sun, Yuhang Cao, Jianze Liang, Qiushi Sun, Ziyu Liu, Zhixiong Zhang, Yuhang Zang, Xiaoyi Dong, Kai Chen, Dahua Lin, Jiaqi Wang:
CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning. CoRR abs/2508.20096 (2025)
[i52]Yibin Wang, Zhimin Li, Yuhang Zang, Yujie Zhou, Jiazi Bu, Chunyu Wang, Qinglin Lu, Cheng Jin, Jiaqi Wang:
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning. CoRR abs/2508.20751 (2025)
[i51]Xilin Wei, Xiaoran Liu, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Jiaqi Wang, Xipeng Qiu, Dahua Lin:
SIM-CoT: Supervised Implicit Chain-of-Thought. CoRR abs/2509.20317 (2025)
[i50]Junbo Niu, Zheng Liu, Zhuangcheng Gu, Bin Wang, Linke Ouyang, Zhiyuan Zhao, Tao Chu, Tianyao He, Fan Wu, Qintong Zhang, Zhenjiang Jin, Guang Liang, Rui Zhang, Wenzheng Zhang, Yuan Qu, Zhifei Ren, Yuefeng Sun, Yuanhong Zheng, Dongsheng Ma, Zirui Tang, Boyu Niu, Ziyang Miao, Hejun Dong, Siyi Qian, Junyuan Zhang, Jingzhou Chen, Fangdong Wang, Xiaomeng Zhao
, Liqun Wei, Wei Li, Shasha Wang, Ruiliang Xu, Yuanyuan Cao, Lu Chen, Qianqian Wu, Huaiyu Gu, Lindong Lu, Keming Wang, Dechen Lin, Guanlin Shen, Xuanhe Zhou, Linfeng Zhang, Yuhang Zang, Xiaoyi Dong, Jiaqi Wang, Bo Zhang, Lei Bai, Pei Chu, Weijia Li, Jiang Wu, Lijun Wu, Zhenxiang Li, Guangyu Wang, Zhongying Tu, Chao Xu, Kai Chen, Yu Qiao, Bowen Zhou, Dahua Lin, Wentao Zhang, Conghui He:
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing. CoRR abs/2509.22186 (2025)
[i49]Ziyu Liu, Yuhang Zang, Shengyuan Ding, Yuhang Cao, Xiaoyi Dong, Haodong Duan, Dahua Lin, Jiaqi Wang:
SPARK: Synergistic Policy And Reward Co-Evolving Framework. CoRR abs/2509.22624 (2025)
[i48]Long Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Jianze Liang, Qidong Huang, Jiaqi Wang, Feng Wu, Dahua Lin:
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning. CoRR abs/2509.22647 (2025)
[i47]Zhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Jiaqi Wang:
2nd Place Report of MOSEv2 Challenge 2025: Concept Guided Video Object Segmentation via SeC. CoRR abs/2509.23838 (2025)
[i46]Yujie Zhou, Pengyang Ling, Jiazi Bu, Yibin Wang, Yuhang Zang, Jiaqi Wang, Li Niu, Guangtao Zhai:
G2RPO: Granular GRPO for Precise Reward in Flow Models. CoRR abs/2510.01982 (2025)
[i45]Chang Liu, Henghui Ding, Kaining Ying, Lingyi Hong, Ning Xu, Linjie Yang, Yuchen Fan, Mingqi Gao, Jingkun Chen, Yunqi Miao, Gengshen Wu, Zhijin Qin, Jungong Han, Zhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Jiaqi Wang, Chang Soo Lim, Joonyoung Moon, Donghyeon Cho, Tingmin Li, Yixuan Li, Yang Yang, An Yan, Leilei Cao, Feng Lu, Ran Hong, Youhai Jiang, Fengjie Zhu, Yujie Xie, Hongyang Zhang, Zhihui Liu, Shihai Ruan, Quanzhu Niu, Dengxian Gong, Shihao Chen, Tao Zhang, Yikang Zhou, Haobo Yuan, Lu Qi, Xiangtai Li, Shunping Ji, Alexey Nekrasov, Ali Athar, Daan de Geus, Alexander Hermans, Bastian Leibe:
LSVOS 2025 Challenge Report: Recent Advances in Complex Video Object Segmentation. CoRR abs/2510.11063 (2025)
[i44]Yibin Wang, Zhimin Li, Yuhang Zang, Jiazi Bu, Yujie Zhou, Yi Xin, Junjun He, Chunyu Wang, Qinglin Lu, Cheng Jin, Jiaqi Wang:
UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation. CoRR abs/2510.18701 (2025)
[i43]Zihan Liu, Zhikang Niu, Qiuyang Xiao, Zhisheng Zheng, Ruoqi Yuan, Yuhang Zang, Yuhang Cao, Xiaoyi Dong, Jianze Liang, Xie Chen, Leilei Sun, Dahua Lin, Jiaqi Wang:
STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence. CoRR abs/2510.24693 (2025)
[i42]Yuhong Liu, Beichen Zhang, Yuhang Zang, Yuhang Cao, Long Xing, Xiaoyi Dong, Haodong Duan, Dahua Lin, Jiaqi Wang:
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning. CoRR abs/2510.27606 (2025)
[i41]Huiqiang Sun, Liao Shen, Zhan Peng, Kun Wang, Size Wu, Yuhang Zang, Tianqi Liu, Zihao Huang, Xingyu Zeng, Zhiguo Cao, Wei Li, Chen Change Loy:
Generative Photographic Control for Scene-Consistent Video Cinematic Editing. CoRR abs/2511.12921 (2025)
[i40]Beichen Zhang, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang:
Think Visually, Reason Textually: Vision-Language Synergy in ARC. CoRR abs/2511.15703 (2025)
[i39]Junyuan Zhang, Bin Wang, Qintong Zhang, Fan Wu, Zichen Wen, Jialin Lu, Junjie Shan, Ziqi Zhao, Shuya Yang, Ziling Wang, Ziyang Miao, Huaping Zhong, Yuhang Zang, Xiaoyi Dong, Ka-Ho Chow, Conghui He:
TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition. CoRR abs/2512.01248 (2025)
[i38]Shengyuan Ding, Xinyu Fang, Ziyu Liu, Yuhang Zang, Yuhang Cao, Xiangyu Zhao, Haodong Duan, Xiaoyi Dong, Jianze Liang, Bin Wang, Conghui He, Dahua Lin, Jiaqi Wang:
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning. CoRR abs/2512.05111 (2025)- 2024
[c17]Zeyi Sun, Ye Fang, Tong Wu, Pan Zhang, Yuhang Zang
, Shu Kong, Yuanjun Xiong, Dahua Lin, Jiaqi Wang
:
Alpha-CLIP: A CLIP Model Focusing on Wherever you Want. CVPR 2024: 13019-13029
[c16]Tianqi Liu
, Guangcong Wang
, Shoukang Hu
, Liao Shen
, Xinyi Ye
, Yuhang Zang
, Zhiguo Cao
, Wei Li, Ziwei Liu
:
MVSGaussian: Fast Generalizable Gaussian Splatting Reconstruction from Multi-View Stereo. ECCV (18) 2024: 37-53
[c15]Beichen Zhang, Pan Zhang, Xiaoyi Dong, Yuhang Zang, Jiaqi Wang
:
Long-CLIP: Unlocking the Long-Text Capability of CLIP. ECCV (51) 2024: 310-325
[c14]Yuhang Zang, Hanlin Goh, Joshua M. Susskind, Chen Huang:
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization. ICLR 2024
[c13]Haodong Duan
, Junming Yang
, Yuxuan Qiao
, Xinyu Fang
, Lin Chen
, Yuan Liu
, Xiaoyi Dong
, Yuhang Zang
, Pan Zhang
, Jiaqi Wang
, Dahua Lin
, Kai Chen
:
VLMEvalKit: An Open-Source ToolKit for Evaluating Large Multi-Modality Models. ACM Multimedia 2024: 11198-11201
[c12]Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Lin Bin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao, Jiaqi Wang:
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions. NeurIPS 2024
[c11]Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, Feng Zhao:
Are We on the Right Way for Evaluating Large Vision-Language Models? NeurIPS 2024
[c10]Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Zhe Chen, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Kai Chen, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao, Dahua Lin, Jiaqi Wang:
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD. NeurIPS 2024
[c9]Ziyu Liu, Tao Chu, Yuhang Zang, Xilin Wei, Xiaoyi Dong, Pan Zhang, Zijian Liang, Yuanjun Xiong, Yu Qiao, Dahua Lin, Jiaqi Wang:
MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs. NeurIPS 2024
[c8]Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, Pan Zhang, Liangming Pan, Yu-Gang Jiang, Jiaqi Wang, Yixin Cao, Aixin Sun:
MMLONGBENCH-DOC: Benchmarking Long-context Document Understanding with Visualizations. NeurIPS 2024
[c7]Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, Jiaqi Wang:
Streaming Long Video Understanding with Large Language Models. NeurIPS 2024
[i37]Yuhang Zang, Hanlin Goh, Josh M. Susskind, Chen Huang:
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization. CoRR abs/2401.15914 (2024)
[i36]Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Xilin Wei, Songyang Zhang
, Haodong Duan, Maosong Cao, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin, Jiaqi Wang:
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model. CoRR abs/2401.16420 (2024)
[i35]Ziyu Liu, Zeyi Sun, Yuhang Zang, Wei Li, Pan Zhang, Xiaoyi Dong, Yuanjun Xiong, Dahua Lin, Jiaqi Wang:
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition. CoRR abs/2403.13805 (2024)
[i34]Beichen Zhang, Pan Zhang, Xiaoyi Dong, Yuhang Zang, Jiaqi Wang
:
Long-CLIP: Unlocking the Long-Text Capability of CLIP. CoRR abs/2403.15378 (2024)
[i33]Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang
, Yu Qiao, Dahua Lin, Feng Zhao:
Are We on the Right Way for Evaluating Large Vision-Language Models? CoRR abs/2403.20330 (2024)
[i32]Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang
, Haodong Duan, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Zhe Chen, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Kai Chen, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao, Dahua Lin, Jiaqi Wang
:
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD. CoRR abs/2404.06512 (2024)
[i31]Tao Chu, Pan Zhang, Xiaoyi Dong, Yuhang Zang, Qiong Liu, Jiaqi Wang:
Unified Scene Representation and Reconstruction for 3D Large Language Models. CoRR abs/2404.13044 (2024)
[i30]Tianqi Liu
, Guangcong Wang, Shoukang Hu, Li Shen, Xinyi Ye, Yuhang Zang, Zhiguo Cao, Wei Li, Ziwei Liu:
Fast Generalizable Gaussian Splatting Reconstruction from Multi-View Stereo. CoRR abs/2405.12218 (2024)
[i29]Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, Jiaqi Wang
:
Streaming Long Video Understanding with Large Language Models. CoRR abs/2405.16009 (2024)
[i28]Zeyi Sun, Tong Wu, Pan Zhang, Yuhang Zang, Xiaoyi Dong, Yuanjun Xiong, Dahua Lin, Jiaqi Wang
:
Bootstrap3D: Improving 3D Content Creation with Synthetic Data. CoRR abs/2406.00093 (2024)
[i27]Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao, Jiaqi Wang
:
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions. CoRR abs/2406.04325 (2024)
[i26]Pengyang Ling, Jiazi Bu, Pan Zhang, Xiaoyi Dong, Yuhang Zang, Tong Wu, Huaian Chen, Jiaqi Wang
, Yi Jin:
MotionClone: Training-Free Motion Cloning for Controllable Video Generation. CoRR abs/2406.05338 (2024)
[i25]Jiaqi Wang, Yuhang Zang, Pan Zhang, Tao Chu, Yuhang Cao, Zeyi Sun, Ziyu Liu, Xiaoyi Dong, Tong Wu, Dahua Lin, Zeming Chen, Zhi Wang, Lingchen Meng, Wenhao Yao, Jianwei Yang, Sihong Wu, Zhineng Chen, Zuxuan Wu, Yu-Gang Jiang, Peixi Wu, Bosong Chai, Xuan Nie, Longquan Yan, Zeyu Wang, Qifan Zhou, Boning Wang, Jiaqi Huang, Zunnan Xu, Xiu Li, Kehong Yuan, Yanyan Zu, Jiayao Ha, Qiong Gao, Licheng Jiao:
V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results. CoRR abs/2406.11739 (2024)
[i24]Ziyu Liu, Tao Chu, Yuhang Zang, Xilin Wei, Xiaoyi Dong, Pan Zhang, Zijian Liang, Yuanjun Xiong, Yu Qiao, Dahua Lin, Jiaqi Wang
:
MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs. CoRR abs/2406.11833 (2024)
[i23]Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, Pan Zhang, Liangming Pan, Yu-Gang Jiang, Jiaqi Wang
, Yixin Cao, Aixin Sun:
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations. CoRR abs/2407.01523 (2024)
[i22]Zihao Huang, Shoukang Hu, Guangcong Wang, Tianqi Liu, Yuhang Zang, Zhiguo Cao, Wei Li, Ziwei Liu:
WildAvatar: Web-scale In-the-wild Video Dataset for 3D Avatar Creation. CoRR abs/2407.02165 (2024)
[i21]Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian, Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang, Songyang Zhang
, Wenwei Zhang, Yining Li, Yang Gao, Peng Sun, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Hang Yan, Conghui He, Xingcheng Zhang, Kai Chen, Jifeng Dai, Yu Qiao, Dahua Lin, Jiaqi Wang:
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output. CoRR abs/2407.03320 (2024)
[i20]Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Jiaqi Wang, Dahua Lin, Kai Chen:
VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models. CoRR abs/2407.11691 (2024)
[i19]Jiazi Bu, Pengyang Ling, Pan Zhang, Tong Wu, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang
:
BroadWay: Boost Your Text-to-Video Generation Model in a Training-free Way. CoRR abs/2410.06241 (2024)
[i18]Qidong Huang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Jiaqi Wang
, Dahua Lin, Weiming Zhang, Nenghai Yu:
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate. CoRR abs/2410.07167 (2024)
[i17]Shuangrui Ding, Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Yuwei Guo, Dahua Lin, Jiaqi Wang
:
SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree. CoRR abs/2410.16268 (2024)
[i16]Long Xing, Qidong Huang, Xiaoyi Dong, Jiajie Lu, Pan Zhang, Yuhang Zang, Yuhang Cao, Conghui He, Jiaqi Wang
, Feng Wu, Dahua Lin:
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction. CoRR abs/2410.17247 (2024)
[i15]Ziyu Liu, Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Haodong Duan, Conghui He, Yuanjun Xiong, Dahua Lin, Jiaqi Wang
:
MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models. CoRR abs/2410.17637 (2024)
[i14]Zeyi Sun, Ziyang Chu, Pan Zhang, Tong Wu, Xiaoyi Dong, Yuhang Zang, Yuanjun Xiong, Dahua Lin, Jiaqi Wang
:
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models. CoRR abs/2412.01824 (2024)
[i13]Pan Zhang, Xiaoyi Dong, Yuhang Cao, Yuhang Zang, Rui Qian, Xilin Wei, Lin Chen, Yifei Li, Junbo Niu, Shuangrui Ding, Qipeng Guo, Haodong Duan, Xin Chen, Han Lv, Zheng Nie, Min Zhang, Bin Wang, Wenwei Zhang, Xinyue Zhang, Jiaye Ge, Wei Li, Jingwen Li, Zhongying Tu, Conghui He, Xingcheng Zhang, Kai Chen, Yu Qiao, Dahua Lin, Jiaqi Wang:
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions. CoRR abs/2412.09596 (2024)- 2023
[j1]Yuhang Zang
, Kaiyang Zhou, Chen Huang, Chen Change Loy:
Semi-Supervised and Long-Tailed Object Detection with CascadeMatch. Int. J. Comput. Vis. 131(4): 987-1001 (2023)
[i12]Yuhang Zang, Kaiyang Zhou, Chen Huang, Chen Change Loy:
Semi-Supervised and Long-Tailed Object Detection with CascadeMatch. CoRR abs/2305.14813 (2023)
[i11]Yuhang Zang, Wei Li, Jun Han, Kaiyang Zhou, Chen Change Loy:
Contextual Object Detection with Multimodal Large Language Models. CoRR abs/2305.18279 (2023)
[i10]Zeyi Sun
, Ye Fang, Tong Wu, Pan Zhang, Yuhang Zang, Shu Kong, Yuanjun Xiong, Dahua Lin, Jiaqi Wang
:
Alpha-CLIP: A CLIP Model Focusing on Wherever You Want. CoRR abs/2312.03818 (2023)- 2022
[c6]Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, Chen Change Loy:
Open-Vocabulary DETR with Conditional Matching. ECCV (9) 2022: 106-122
[i9]Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, Chen Change Loy:
Open-Vocabulary DETR with Conditional Matching. CoRR abs/2203.11876 (2022)
[i8]Kaiyang Zhou, Yuanhan Zhang, Yuhang Zang, Jingkang Yang, Chen Change Loy, Ziwei Liu:
On-Device Domain Generalization. CoRR abs/2209.07521 (2022)
[i7]Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, Chen Change Loy:
Unified Vision and Language Prompt Learning. CoRR abs/2210.07225 (2022)- 2021
[c5]Jiaqi Wang
, Wenwei Zhang, Yuhang Zang, Yuhang Cao, Jiangmiao Pang, Tao Gong, Kai Chen, Ziwei Liu, Chen Change Loy, Dahua Lin:
Seesaw Loss for Long-Tailed Instance Segmentation. CVPR 2021: 9695-9704
[c4]Yuhang Zang, Chen Huang, Chen Change Loy:
FASA: Feature Augmentation and Sampling Adaptation for Long-Tailed Instance Segmentation. ICCV 2021: 3437-3446
[i6]Yuhang Zang, Chen Huang, Chen Change Loy:
FASA: Feature Augmentation and Sampling Adaptation for Long-Tailed Instance Segmentation. CoRR abs/2102.12867 (2021)- 2020
[c3]Guanglu Song, Yu Liu, Yuhang Zang, Xiaogang Wang, Biao Leng, Qingsheng Yuan:
KPNet: Towards Minimal Face Detector. AAAI 2020: 12015-12022
[i5]Guanglu Song, Yu Liu, Yuhang Zang
, Xiaogang Wang, Biao Leng, Qingsheng Yuan:
KPNet: Towards Minimal Face Detector. CoRR abs/2003.07543 (2020)
[i4]Yu Liu, Guanglu Song, Yuhang Zang
, Yan Gao, Enze Xie, Junjie Yan, Chen Change Loy, Xiaogang Wang:
1st Place Solutions for OpenImage2019 - Object Detection and Instance Segmentation. CoRR abs/2003.07557 (2020)
[i3]Jiaqi Wang, Wenwei Zhang, Yuhang Zang, Yuhang Cao, Jiangmiao Pang, Tao Gong, Kai Chen, Ziwei Liu, Chen Change Loy, Dahua Lin:
Seesaw Loss for Long-Tailed Instance Segmentation. CoRR abs/2008.10032 (2020)
2010 – 2019
- 2019
[c2]Enze Xie, Yuhang Zang
, Shuai Shao
, Gang Yu
, Cong Yao, Guangyao Li:
Scene Text Detection with Supervised Pyramid Context Network. AAAI 2019: 9038-9045
[c1]Wenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang
, Wenjia Wang, Tong Lu, Gang Yu
, Chunhua Shen:
Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation Network. ICCV 2019: 8439-8448
[i2]Wenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang, Wenjia Wang, Tong Lu, Gang Yu, Chunhua Shen:
Efficient and Accurate Arbitrary-Shaped Text Detection with Pixel Aggregation Network. CoRR abs/1908.05900 (2019)- 2018
[i1]Enze Xie, Yuhang Zang, Shuai Shao, Gang Yu, Cong Yao, Guangyao Li:
Scene Text Detection with Supervised Pyramid Context Network. CoRR abs/1811.08605 (2018)
Coauthor Index

manage site settings
To protect your privacy, all features that rely on external API calls from your browser are turned off by default. You need to opt-in for them to become active. All settings here will be stored as cookies with your web browser. For more information see our F.A.Q.
Unpaywalled article links
Add open access links from
to the list of external document links (if available).
Privacy notice: By enabling the option above, your browser will contact the API of unpaywall.org to load hyperlinks to open access articles. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Unpaywall privacy policy.
Archived links via Wayback Machine
For web page which are no longer available, try to retrieve content from the
of the Internet Archive (if available).
Privacy notice: By enabling the option above, your browser will contact the API of archive.org to check for archived content of web pages that are no longer available. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Internet Archive privacy policy.
Reference lists
Add a list of references from
,
, and
to record detail pages.
load references from crossref.org and opencitations.net
Privacy notice: By enabling the option above, your browser will contact the APIs of crossref.org, opencitations.net, and semanticscholar.org to load article reference information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Crossref privacy policy and the OpenCitations privacy policy, as well as the AI2 Privacy Policy covering Semantic Scholar.
Citation data
Add a list of citing articles from
and
to record detail pages.
load citations from opencitations.net
Privacy notice: By enabling the option above, your browser will contact the API of opencitations.net and semanticscholar.org to load citation information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the OpenCitations privacy policy as well as the AI2 Privacy Policy covering Semantic Scholar.
OpenAlex data
Load additional information about publications from
.
Privacy notice: By enabling the option above, your browser will contact the API of openalex.org to load additional information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the information given by OpenAlex.
last updated on 2026-01-24 04:04 CET by the dblp team
all metadata released as open data under CC0 1.0 license
see also: Terms of Use | Privacy Policy | Imprint


Google
Google Scholar
Semantic Scholar
Internet Archive Scholar
CiteSeerX
ORCID







