Yiming Tang

Yiming Tang

Email: yiming at nus dot edu dot sg

Google Scholar: View Profile

Curriculum Vitae: View CV

WeChat: Yiming_Tangible

Location: MD1, 12 Science Drive 2, Singapore

About Me

Hi! I'm Yiming, a third-year PhD student in National University of Singapore, very fortunate to be advised by Prof. Dianbo Liu. I lead the interpretability research group in Artificial Scientific Intelligence Lab. My research primarily aims to develop algorithms and theories toward a better understanding of intelligence for both scientific research and AI safety, currently centered around the identification of human-interpretable concepts encoded in LMMs, and shifting toward tracking representation drifts induced by post-training. During my undergraduate studies, I was very lucky to collaborate with Prof. Bin Dong, Prof. Shanghang Zhang, and Prof. Hao Dong.

News

Research Interests

My research interests focus on:

Beyond these topics, I am also broadly interested in LMMs and AI Agents. See What We Work On.

Collaboration

I'm highly open for mutually beneficial collaborations with the central aim to advance AI Interpretability and AI safety via scientific research. Researchers working on aligned research topics are especially welcome. I can also offer research intern positions for undergraduate and master students. Feel free to contact me directly via email or WeChat if you are interested in our research.

Selected Works

A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima

We develop the first theoretical framework for sparse dictionary learning in mechanistic interpretability.

Paper
SDL Theory Diagram
How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders

We develop matryoshka transcoders and utilize it to identify physical plausibility failure modes of generative models, achieving SOTA performances in targeted feature discovery.

Paper
Matryoshka Transcoders Diagram
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders

We develop LanSE and decompose natural and medical images into interpretable visual patterns grounded in natural language, supporting fine-grained analysis on AI generated contents.

Paper
LanSE Diagram
LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model

We are the first in the literature to prompt engineer LLMs to inspect datasets and analyze their subpopulation structures, paving the way for advanced dataset analysis with LLMs.

Paper
LLM as Dataset Analyst Diagram
Prompt Engineering Through the Lens of Optimal Control

We are one of the first approaches to develop a theoretical framework to understand various prompt engineering methods through the lens of rigorous optimal control theory.

Paper
Prompt Optimal Control Diagram
Introducing Brainiac Buddy, your AI-powered teaching assistant

It was a real honor to participate in the early development of Brainiac Buddy, an AI-powered teaching assistant with real-world applications in Peking University developed by Bin Dong's team.

Project
Brainiac Buddy Diagram
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control

We are the first in the literature to apply gradient ascent on MechInterp features to support prompt engineering for persona control.

Paper
Bridging Mechanistic Interpretability and Prompt Engineering Diagram
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

We identify, analyze, and mitigate the existence of textual biases in visual tokens in vision-language models.

Paper
Over-Alignment Diagram
The Integrated Forward-Forward Algorithm: Integrating Forward-Forward and Shallow Backpropagation with Local Losses

We follow Hinton's Forward-Forward algorithm and make it applicable for deeper layers via the introduction of local losses.

Paper
Integrated Forward-Forward Diagram

Publications and Preprints

A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
Paper
How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
Paper
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Paper
LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model
Paper
Prompt Engineering Through the Lens of Optimal Control
Paper
The Integrated Forward-Forward Algorithm: Integrating Forward-Forward and Shallow Backpropagation with Local Losses
Paper
Demonstration Notebook: Finding the Most Suited In-Context Learning Example from Interactions
Paper
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
Paper
CXR-LanIC: Language-Grounded Interpretable Classifier for Chest X-Ray Diagnosis
Paper
Benchmarking Machine Learning Agents for Scientific Research
Paper
SAN: Hypothesizing Long-Term Synaptic Development and Neural Engram Mechanism in Scalable Model's Parameter-Efficient Fine-Tuning
Paper
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
Paper
Rep2Text: Decoding Full Text from a Single LLM Token Representation
Paper

For a complete list of publications, please visit my Google Scholar profile.

Education

NUS Logo
National University of Singapore

Doctor of Philosophy, College of Design and Engineering

Peking University Logo
Peking University

Bachelor of Science, School of Mathematical Sciences

Talks

Mechanistic Interpretability on Neural Representations

National University of Singapore

July 20th, 2026

Mechanistic Interpretability and AI Safety

Alibaba

July 8th, 2026

Sparse Dictionary Learning on fMRI

National University of Singapore

January 28th, 2026

Language-Grounded Sparse Encoders

National University of Singapore

October 15th, 2025

Open Problems in Mechanistic Interpretability - Paper Walkthrough

National University of Singapore

July 10th, 2025

An Introduction to Methods for LLM Explainability

Peking University

April 22th, 2024

An Introduction to LLM as Dataset Analyst and PE Control

Peking University

March 28th, 2024

Yiming's Undergraduate Research Sharing

Peking University

March 11th, 2024

Experiences

Alibaba Logo
Research Intern

Alibaba Security, Alibaba Group Holding Limited, Hangzhou

June 2026 - Present

Supervisor: Hang Zhou

Cognitive AI for Science Lab Logo
Research Associate

Artificial Scientific Intelligence Lab, National University of Singapore, Singapore

August 2024 - Present

Supervisor: Dianbo Liu

HMI Lab Logo
Research Intern

Human Machine Intelligence Lab, Peking University, Beijing

2023 - 2024

Supervisor: Shanghang Zhang

BICMR Logo
Research Intern

Beijing International Center for Mathematical Research, Peking University, Beijing

2023 - 2024

Supervisor: Bin Dong

Hyperplane Lab Logo
Research Intern

Hyperplane Lab, Peking University, Beijing

2021 - 2022

Supervisor: Hao Dong

Aijianzi Logo
Teaching Researcher

Aijianzi, Beijing Aijianzi Education Technology, Beijing

2019 - 2021

Supervisor: Xiaojun Hu

Mentorship

I have had the privilege of mentoring several students working on research topics in Mechanistic Interpretability, Neuroscience, and AI for Biomedical Sciences.

Harshvardhan Saini
Harshvardhan Saini

Indian Institute of Technology, Dhanbad, India

Zheng Lin
Zheng Lin

Hong Kong University of Science and Technology, Hong Kong, China

Zhaoqian Yao
Zhaoqian Yao

Chinese University of Hong Kong, Hong Kong, China

Qinglin Qi
Qinglin Qi

Lund University, Lund, Sweden

Lingheng Du
Lingheng Du

Peking University, Beijing, China

Yang TuαΊ₯n Anh
Yang TuαΊ₯n Anh

National University of Singapore, Singapore

Meet Desmond

An honor and a pleasure to introduce to you here my most adorable cat, Desmond, the docent of my research and the guardian of this website. Have a chat and he will answer your questions in my absence.

Black Cat Cyber Pet
Wellcome, visitor! Master is out, let us talk. 🐱