|
|
Scalable Visual Pretraining for Language Intelligence
Yiming Zhang,
Zhonghan Zhao,
Wenwei Zhang,
Haiteng Zhao,
Tianyang Lin,
Huanze Tang,
Yunhua Zhou,
Demin Song,
Kuikun Liu,
Haochen Ye,
Haian Huang,
Yuzhe Gu,
Haijun Lv,
Qipeng Guo,
Bin Liu,
Gaoang Wang,
Kai Chen
arXiv preprint, 2026
arXiv
A systematic study showing that visual pretraining directly on rendered documents consistently outperforms text-only pretraining on the same corpora, offering an efficient pathway to scalable language intelligence.
|
|
|
FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Yiming Zhang,
Yicheng Gu,
Yanhong Zeng,
Zhening Xing,
Yuancheng Wang,
Zhizheng Wu,
Kai Chen
IJCV, 2026
project page /
video /
arXiv /
demo /
code
FoleyCrafter is a video-to-audio generation framework which can produce realistic sound effects semantically relevant and synchronized with videos.
|
|
|
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models
Yiming Zhang*,
Zhening Xing*,
Yanhong Zeng,
Youqing Fang,
Kai Chen
CVPR, 2024
project page /
video /
arXiv /
demo /
code
PIA can animate any images from personalized models by text while preserving high-fidelity details and unique styles.
|
|
|
Exploring Visual Pretraining for Learning Language Intelligence
Zhonghan Zhao*,
Yiming Zhang*,
Wenwei Zhang,
Haiteng Zhao,
Xingguang Wei,
Zhangwei Gao,
Kuikun Liu,
Yuzhe Gu,
Size Wu,
Haian Huang,
Jianfei Gao,
Haijun Lv,
Demin Song,
Yunhua Zhou,
Qipeng Guo,
Gaoang Wang,
Kai Chen
CVPR, 2026
paper
A first attempt (MAPLE) showing that LLMs can be pretrained on visual images of documents to reach parity with text-only models, breaking the text-scaling bottleneck.
|
|
|
Achieving Olympiad-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
Haiteng Zhao*,
Junhao Shen*,
Yiming Zhang*,
Songyang Gao,
Kuikun Liu,
Tianyou Ma,
Fan Zheng,
Dahua Lin,
Wenwei Zhang,
Kai Chen
ICLR, 2026
arXiv
InternGeometry is an LLM agent achieving medalist-level performance on olympiad geometry via complexity-boosting reinforcement learning, using only 0.04% of the data used by AlphaGeometry 2.
|
|
|
Intern-S2-Preview: Scientific Agentic Foundation Model
Lei Bai, Jiaqi Cao, Chiyu Chen, ..., Yiming Zhang, ..., Kai Chen, et al.
arXiv preprint, 2026
arXiv
Intern-S2-Preview is a series of scientific agentic foundation models supporting multimodal scientific understanding, reasoning, generation, and long-horizon agentic tasks.
|
- Invited Talk: Personalized Image Animator (OpenMMLab on Bilibili Live 2024)
- Conference Reviewer: NeurIPS, CVPR, ICCV, ECCV.
-
Outstanding Undergraduate in 2024.
-
National Scholarship in 2022 (Top 1% in DUT).
- Merit Students.
-
First Prize Excellence Scholarship.
|