Yiwei Chen陈一苇

I am a CS Ph.D. student focusing on Trustworthy ML and Scalable AI at Michigan State University (MSU), where my advisor is Prof. Sijia Liu. Before that, I earned both my Bachelor's and Master's degrees from Xi'an Jiaotong University (XJTU), where I was admitted via the Special Class for the Gifted Young.

profile photo
Research

My research interests lie at the intersection of robust and explainable artificial intelligence, with the long-term goal of making AI systems safe and scalable. Currently, I'm working on:

  • Foundation Models: Large Language Models, Multi-Modal Language Models, Multi-Modality
  • Agent & Post-Training: Agentic Systems, GUI Agents, Reasoning, Post-Training
  • Trustworthy AI: Alignment, Trustworthy Algorithms, Interpretability, Machine Unlearning
News
  • 2026-07 Two papers accepted at COLM'26, including one as first author!
  • 2026-05 Joined Amazon as an Applied Scientist Intern this summer!
  • 2026-01 Two first-author papers accepted at ICLR'26!
  • 2025-05 Joined Cisco as a PhD Intern!
  • 2025-05 Unlearning Isn't Invisible on detecting unlearning traces in LLMs from model outputs!
  • 2025-03 Safety Mirage on the application of machine unlearning to VLM safety alignment, my first project as a PhD student!
  • 2024-08 Started PhD journey at Michigan State University!
  • 2024-06 Graduated from Xi'an Jiaotong University!
Internships
Amazon Applied Scientist Intern May 2026 – Aug. 2026
Cisco PhD Intern May 2025 – Oct. 2025
Selected Publications

(* denotes equal contribution)

VLM Unlearn
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
Yiwei Chen*, Yuguang Yao*, Yihua Zhang, Bingquan Shen, Gaowen Liu, Sijia Liu
ICLR 2026
CodePaper

Unlearning removes the harm that safety fine-tuning only masks.

Safety fine-tuning of MLLMs inherits the bias of its training data, producing spurious correlations and over-rejections under one-word attacks.

Unlearn Trace
Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs
Yiwei Chen*, Soumyadeep Pal*, Yimeng Zhang, Qing Qu, Sijia Liu
ICLR 2026Oral @ MUGen ICML'25
CodePaper

Forgetting is not invisible: unlearned models still leak that they were unlearned.

Unlearned LLMs keep persistent fingerprints in their outputs and hidden activations, enough for a classifier to tell that unlearning happened at all.

LLM Lineage
Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space
Yiwei Chen*, Bingqi Shang*, Sijia Liu
COLM 2026
ProjectCodePaper

Model provenance from weights alone, with no access to training data.

A geometric fingerprint of weight space traces which model a checkpoint descends from, using spectral energy for coarse discrimination and subspace alignment for fine-grained analysis.

Cybersecurity Exploit Benchmark
Data-Centric Benchmarking of Exploit Generation in LLMs: Understanding the Impact of Fine-Tuning
Yiwei Chen*, Lichi Li*, Kai Cheung, Vinny Parla, Ganesh Sundaram
Technical Report
Paper

With the right fine-tuning data, an 8B model matches frontier LLMs at exploit generation.

The first data-centric benchmark for CVE-conditioned exploit generation: a 6-level context hierarchy, an 8-criterion evaluation, and 17 LLMs measured against it.

Backdoor Unlearn
Forgetting to Forget: Attention Sink as a Gateway for Backdooring LLM Unlearning
Bingqi Shang*, Yiwei Chen*, Yihua Zhang, Bingquan Shen, Sijia Liu
COLM 2026
CodePaper

Attention sinks are a backdoor that quietly restores “forgotten” knowledge.

Triggers planted at attention sinks let an unlearned model recover the forgotten knowledge on cue, while behaving normally whenever the trigger is absent.

MFTR
Tile Classification Based Viewport Prediction with Multi-modal Fusion Transformer
Zhihao Zhang*, Yiwei Chen*, Weizhan Zhang, Caixia Yan, Qinghua Zheng, Qi Wang, Wangdu Chen
ACM MM 2023
CodePaper

Framing viewport prediction as tile classification is what buys the robustness.

Viewport prediction recast as tile classification, using a multi-modal fusion transformer.