2026

AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality Missing Prompt Tuning
AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality Missing Prompt Tuning

Jian Lang, Rongpei Hong, Ting Zhong, Fan Zhou# (# corresponding author)

International Conference on Machine Learning (ICML) 2026

Modality-missing prompt tuning for Multimodal Transformers can unintentionally restrict reasoning to observed-modality subspaces. AOEPT introduces modal-contextualized prompts that recover missing-modality information sources with minimal overhead...

AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality Missing Prompt Tuning

Jian Lang, Rongpei Hong, Ting Zhong, Fan Zhou# (# corresponding author)

International Conference on Machine Learning (ICML) 2026

Modality-missing prompt tuning for Multimodal Transformers can unintentionally restrict reasoning to observed-modality subspaces. AOEPT introduces modal-contextualized prompts that recover missing-modality information sources with minimal overhead...

LEAF: Towards Lightweight Explainable Hateful Video Detection via Self-Grounding CoT Guided Stage-Wise Distillation
LEAF: Towards Lightweight Explainable Hateful Video Detection via Self-Grounding CoT Guided Stage-Wise Distillation

Jian Lang, Rongpei Hong, Meihui Zhong, Kaiju Li, Ting Zhong, Qiang Gao, Fan Zhou# (# corresponding author)

Findings of the Association for Computational Linguistics (ACL Finding) 2026

Hateful video detection remains hard to trust because existing systems are often opaque, while LMM explanations are costly and biased toward benign predictions. LEAF distills self-grounding CoT explanations from LMMs into lightweight SMMs for accurate and interpretable HVD...

LEAF: Towards Lightweight Explainable Hateful Video Detection via Self-Grounding CoT Guided Stage-Wise Distillation

Jian Lang, Rongpei Hong, Meihui Zhong, Kaiju Li, Ting Zhong, Qiang Gao, Fan Zhou# (# corresponding author)

Findings of the Association for Computational Linguistics (ACL Finding) 2026

Hateful video detection remains hard to trust because existing systems are often opaque, while LMM explanations are costly and biased toward benign predictions. LEAF distills self-grounding CoT explanations from LMMs into lightweight SMMs for accurate and interpretable HVD...

MATCH: Multi-Agentic Evidence Grounding for Explainable Hate Video Detection
MATCH: Multi-Agentic Evidence Grounding for Explainable Hate Video Detection

Kaiju Li, Rongpei Hong, Jian Lang, Jin Wu, Fan Zhou, Jingkuan Song

IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) 2026

The growing prevalence of hate videos promoting intolerance, bigotry, and discrimination presents significant psychosocial threats to both individuals and society. Current detection methods often rely on black-box models, which lack interpretability — a crucial factor for fostering more reliable content moderation and trustworthy AI ...

MATCH: Multi-Agentic Evidence Grounding for Explainable Hate Video Detection

Kaiju Li, Rongpei Hong, Jian Lang, Jin Wu, Fan Zhou, Jingkuan Song

IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) 2026

The growing prevalence of hate videos promoting intolerance, bigotry, and discrimination presents significant psychosocial threats to both individuals and society. Current detection methods often rely on black-box models, which lack interpretability — a crucial factor for fostering more reliable content moderation and trustworthy AI ...

2025

TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized Assistant
TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized Assistant

Rongpei Hong, Jian Lang, Ting Zhong, Fan Zhou

ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) 2026

Multimodal Large Language Model (MLLM) Personalization is a critical research problem that facilitates personalized dialogues with MLLMs targeting specific entities (known as personalized concepts). However, existing methods and benchmarks focus on ...

TAMEing Long Contexts in Personalization: Towards Training-Free and State-Aware MLLM Personalized Assistant

Rongpei Hong, Jian Lang, Ting Zhong, Fan Zhou

ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) 2026

Multimodal Large Language Model (MLLM) Personalization is a critical research problem that facilitates personalized dialogues with MLLMs targeting specific entities (known as personalized concepts). However, existing methods and benchmarks focus on ...

Nipping Rumors in the Bud: Retrieval-Guided Topic Adaptation for Test-Time Detection of Fake News Videos
Nipping Rumors in the Bud: Retrieval-Guided Topic Adaptation for Test-Time Detection of Fake News Videos

Jian Lang, Rongpei Hong, Ting Zhong, Yong Wang, Fan Zhou

ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) 2026

Fake News Video Detection is critical for social stability. Existing methods typically assume consistent news topic distribution between training and test phases, failing to detect fake news videos tied to emerging events and unseen topics. To bridge this gap, we introduce ...

Nipping Rumors in the Bud: Retrieval-Guided Topic Adaptation for Test-Time Detection of Fake News Videos

Jian Lang, Rongpei Hong, Ting Zhong, Yong Wang, Fan Zhou

ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) 2026

Fake News Video Detection is critical for social stability. Existing methods typically assume consistent news topic distribution between training and test phases, failing to detect fake news videos tied to emerging events and unseen topics. To bridge this gap, we introduce ...

From Shallow Humor to Metaphor: Towards Label-Free Harmful Meme Detection via LMM Agent Self-Improvement
From Shallow Humor to Metaphor: Towards Label-Free Harmful Meme Detection via LMM Agent Self-Improvement

Jian Lang, Rongpei Hong, Ting Zhong, Leiting Chen, Qiang Gao, Fan Zhou

ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) 2026

The proliferation of harmful memes on social media poses significant risks to public health and social stability. Existing detection methods heavily rely on large-scale labeled data for training, which necessitates substantial manual annotation efforts and limits their adaptability to the continually evolving nature of harmful content.

From Shallow Humor to Metaphor: Towards Label-Free Harmful Meme Detection via LMM Agent Self-Improvement

Jian Lang, Rongpei Hong, Ting Zhong, Leiting Chen, Qiang Gao, Fan Zhou

ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) 2026

The proliferation of harmful memes on social media poses significant risks to public health and social stability. Existing detection methods heavily rely on large-scale labeled data for training, which necessitates substantial manual annotation efforts and limits their adaptability to the continually evolving nature of harmful content.

Borrowing Eyes for the Blind Spot: Overcoming Data Scarcity in Malicious Video Detection via Cross-Domain Retrieval Augmentation
Borrowing Eyes for the Blind Spot: Overcoming Data Scarcity in Malicious Video Detection via Cross-Domain Retrieval Augmentation

Rongpei Hong*, Jian Lang*, Ting Zhong, Fan Zhou (* equal contribution)

International Conference on Computer Vision (ICCV) 2025

The rapid proliferation of online video-sharing platforms has accelerated the spread of malicious videos, creating an urgent need for robust detection methods. However, the performance and generalizability of existing detection approaches are severely limited ...

Borrowing Eyes for the Blind Spot: Overcoming Data Scarcity in Malicious Video Detection via Cross-Domain Retrieval Augmentation

Rongpei Hong*, Jian Lang*, Ting Zhong, Fan Zhou (* equal contribution)

International Conference on Computer Vision (ICCV) 2025

The rapid proliferation of online video-sharing platforms has accelerated the spread of malicious videos, creating an urgent need for robust detection methods. However, the performance and generalizability of existing detection approaches are severely limited ...

Enhancing Fake News Video Detection with Self-Driven Question-Answer from LMMs
Enhancing Fake News Video Detection with Self-Driven Question-Answer from LMMs

Fang Liu, Yili Li, Jian Lang, Rongpei Hong#, Fan Zhou (# corresponding author)

Information Processing & Management (IPM) 2025

The widespread dissemination of fake news on online video sharing platforms endangers the politics and public health. Existing approaches to Fake News Video Detection primarily focus on ...

Enhancing Fake News Video Detection with Self-Driven Question-Answer from LMMs

Fang Liu, Yili Li, Jian Lang, Rongpei Hong#, Fan Zhou (# corresponding author)

Information Processing & Management (IPM) 2025

The widespread dissemination of fake news on online video sharing platforms endangers the politics and public health. Existing approaches to Fake News Video Detection primarily focus on ...

REDEEMing Modality Information Loss: Retrieval-Guided Conditional Generation for Severely Modality Missing Learning
REDEEMing Modality Information Loss: Retrieval-Guided Conditional Generation for Severely Modality Missing Learning

Jian Lang, Rongpei Hong, Zhangtao Cheng, Ting Zhong, Yong Wang, Fan Zhou

ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) 2025

Traditional multimodal learning approaches often assume that all modalities are available during both the training and inference phases. However, this assumption is often impractical in real-world ...

REDEEMing Modality Information Loss: Retrieval-Guided Conditional Generation for Severely Modality Missing Learning

Jian Lang, Rongpei Hong, Zhangtao Cheng, Ting Zhong, Yong Wang, Fan Zhou

ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) 2025

Traditional multimodal learning approaches often assume that all modalities are available during both the training and inference phases. However, this assumption is often impractical in real-world ...

REAL: Retrieval-Augmented Prototype Alignment for Improved Fake News Video Detection
REAL: Retrieval-Augmented Prototype Alignment for Improved Fake News Video Detection

Yili Li, Jian Lang, Rongpei Hong, Qing Chen, Zhangtao Cheng, Jia Chen, Ting Zhong, Fan Zhou

IEEE International Conference on Multimedia and Expo (ICME) 2025

Detecting fake news videos has emerged as a critical task due to their profound implications in politics, finance, and public health. However, existing methods often fail to ...

REAL: Retrieval-Augmented Prototype Alignment for Improved Fake News Video Detection

Yili Li, Jian Lang, Rongpei Hong, Qing Chen, Zhangtao Cheng, Jia Chen, Ting Zhong, Fan Zhou

IEEE International Conference on Multimedia and Expo (ICME) 2025

Detecting fake news videos has emerged as a critical task due to their profound implications in politics, finance, and public health. However, existing methods often fail to ...

Following Clues, Approaching the Truth: Explainable Micro-Video Rumor Detection via Chain-of-Thought Reasoning
Following Clues, Approaching the Truth: Explainable Micro-Video Rumor Detection via Chain-of-Thought Reasoning

Rongpei Hong, Jian Lang, Jin Xu, Zhangtao Cheng, Ting Zhong, Fan Zhou

The ACM Web Conference (WWW) 2025

The rapid spread of rumor content on online micro-video platforms poses significant threats to public health and safety. However, existing Micro-Video Rumor Detection (MVRD) methods are generally black-box, which lacks transparency and makes it difficult to understand the reasoning behind classification decisions...

Following Clues, Approaching the Truth: Explainable Micro-Video Rumor Detection via Chain-of-Thought Reasoning

Rongpei Hong, Jian Lang, Jin Xu, Zhangtao Cheng, Ting Zhong, Fan Zhou

The ACM Web Conference (WWW) 2025

The rapid spread of rumor content on online micro-video platforms poses significant threats to public health and safety. However, existing Micro-Video Rumor Detection (MVRD) methods are generally black-box, which lacks transparency and makes it difficult to understand the reasoning behind classification decisions...

Biting Off More Than You Can Detect: Retrieval-Augmented Multimodal Experts for Short Video Hate Detection
Biting Off More Than You Can Detect: Retrieval-Augmented Multimodal Experts for Short Video Hate Detection

Jian Lang, Rongpei Hong, Jin Xu, Yili Li, Xovee Xu, Fan Zhou

The ACM Web Conference (WWW) 2025

Short Video Hate Detection (SVHD) is increasingly vital as hateful content — such as racial and gender-based discrimination — spreads rapidly across platforms like TikTok, YouTube Shorts, and Instagram Reels. Existing approaches face significant challenges:...

Biting Off More Than You Can Detect: Retrieval-Augmented Multimodal Experts for Short Video Hate Detection

Jian Lang, Rongpei Hong, Jin Xu, Yili Li, Xovee Xu, Fan Zhou

The ACM Web Conference (WWW) 2025

Short Video Hate Detection (SVHD) is increasingly vital as hateful content — such as racial and gender-based discrimination — spreads rapidly across platforms like TikTok, YouTube Shorts, and Instagram Reels. Existing approaches face significant challenges:...

2023

Simplifying Temporal Heterogeneous Network for Continuous-Time Link Prediction
Simplifying Temporal Heterogeneous Network for Continuous-Time Link Prediction

Ce Li, Rongpei Hong, Xovee Xu, Goce Trajcevski, Fan Zhou

The ACM Conference on Information and Knowledge Management (CIKM) 2023

Temporal heterogeneous networks (THNs) investigate the structural interactions and their evolution over time in graphs with multiple types of nodes or edges. Existing THNs describe evolving networks as a sequence of graph snapshots and adopt mechanisms from static heterogeneous networks to capture the spatial-temporal correlation...

Simplifying Temporal Heterogeneous Network for Continuous-Time Link Prediction

Ce Li, Rongpei Hong, Xovee Xu, Goce Trajcevski, Fan Zhou

The ACM Conference on Information and Knowledge Management (CIKM) 2023

Temporal heterogeneous networks (THNs) investigate the structural interactions and their evolution over time in graphs with multiple types of nodes or edges. Existing THNs describe evolving networks as a sequence of graph snapshots and adopt mechanisms from static heterogeneous networks to capture the spatial-temporal correlation...