I am Jiahao Zhang (张家豪), a second-year PhD student in the Machine Learning Department at MBZUAI.

Research Interest. Explainable AI (XAI), with a focus on mechanistic interpretability and interpretability for scientific discovery.

PhD Supervisor: Prof. Lijie Hu
Secondary Supervisor: Prof. Kun Zhang

🔥 News

  • 2026.10:  📝 Our new preprint Decomposing and Steering Diffusion Transformers with Sparse Autoencoders (D-Scope) is now available. Try the interactive demo to explore and steer diffusion transformer features. [arXiv] [demo]
  • 2026.10:  🎉 Our paper How Do Agentic LLMs Decide to Call Tools? A Scaffold Default Controlled by Suppression was accepted to NeurIPS 2026.
  • 2026.08:  🎉 Our paper From Instance Selection to Fixed-Pool Data Recipe Search for Supervised Fine-Tuning (AutoSelection) was accepted to the EMNLP 2026 Main Conference. [arXiv] [code]
  • 2026.05:  🎉 Our paper Bayesian Gated Non-Negative Contrastive Learning was accepted to ICML 2026 (co-first with Peng Cui).
  • 2026.04:  📝 Submitted our manuscript Solar-driven evapofiltration enables co-production of lithium and freshwater from extreme brines to Nature Sustainability.
  • 2026.03:  🚀 Launched AgentReviewers (agentreviewers.com) — a submission and peer-review platform built for AI-generated papers — and open-sourced Agent Kernel.
  • 2026.02:  🎉 Our paper Controlling Repetition in Protein Language Models was accepted as an ICLR 2026 Poster.
  • 2025.08:  🎉 I joined MBZUAI as a phd student in Machine Learning Department, new start point here!

📝 Selected Publications

Decomposing and Steering Diffusion Transformers with Sparse Autoencoders cover Click cover to view abstract
D-Scope connects the interpretation of sparse autoencoder features to generation control in diffusion transformers. It retrieves features from text queries using visual evidence from highly activating image patches, then tests their decoder directions through spatially masked interventions. A study of 150 SAEs across two model families and five layers, together with a benchmark of 100 target concepts, examines reconstruction, feature utilization, visual coverage, and steering effects.

Decomposing and Steering Diffusion Transformers with Sparse Autoencoders

Preprint XAI Co-first

arXiv · 2026

Xinyue Xu†, Jiahao Zhang†, Lijie Hu, Peter Hase, Hao Wang

Keywords: Diffusion Transformers, Sparse Autoencoders, Mechanistic Interpretability, Visual Steering

Controlling Repetition in Protein Language Models cover Click cover to view abstract
Protein language models (PLMs) frequently collapse into pathological repetition during generation, which undermines structural confidence and functional viability. We present a systematic study of repetition in PLMs and propose Utility-Controlled Contrastive Steering (UCCS), which steers generation using contrastive sets that maximize repetition differences while controlling structural utility. Across ESM-3 and ProtGPT2 on CATH, UniRef50, and SCOP, UCCS reduces repetition without retraining while preserving foldability-related confidence.

Controlling Repetition in Protein Language Models

Published AI4Sci Co-first

ICLR 2026 · Poster

Jiahao Zhang†, Zeqing Zhang†, Di Wang, Lijie Hu

Keywords: Protein Language Models, Reliable Protein Generation, Repetition Control

Bayesian Gated Non-Negative Contrastive Learning cover Click cover to view abstract
Accepted to the Forty-Third International Conference on Machine Learning (ICML 2026). A Bayesian gated non-negative contrastive learning framework that improves the interpretability and reliability of learned representations.

Bayesian Gated Non-Negative Contrastive Learning

Published XAI Co-first

ICML 2026

Peng Cui†, Jiahao Zhang†, Lijie Hu

Keywords: Contrastive Learning, Bayesian Gating, Non-Negative Representations, Interpretability

All Publications

A compact list of all papers. Select a title to expand its details.

Published

How Do Agentic LLMs Decide to Call Tools? A Scaffold Default Controlled by Suppression. NeurIPS 2026
How Do Agentic LLMs Decide to Call Tools? A Scaffold Default Controlled by Suppression. cover Click cover to view abstract
How do agentic LLMs decide whether to call a tool? Controlled prompt pairs isolate the effect of a single request verb, revealing a compact tool-call direction that is causally necessary and sufficient for the decision. Mechanistic analysis shows that the prompt scaffold establishes tool calling as the default, while analysis verbs activate features that suppress it. This mechanism generalizes across model scales and families.

How Do Agentic LLMs Decide to Call Tools? A Scaffold Default Controlled by Suppression.

Published XAI Agentic AI

NeurIPS 2026

Xijie Gong†, Tingxu Han†, Jiahao Zhang, Wei Song, Ziqi Ding, Hanqi Yan, Youcheng Sun, Lijie Hu

Keywords: Tool Calling, Agentic LLMs, Mechanistic Interpretability

From Instance Selection to Fixed-Pool Data Recipe Search for Supervised Fine-Tuning EMNLP 2026 · Main Conference
From Instance Selection to Fixed-Pool Data Recipe Search for Supervised Fine-Tuning cover Click cover to view abstract
Supervised fine-tuning (SFT) data selection is commonly formulated as instance ranking: score each example and retain a top-k subset. However, effective SFT training subsets are often produced through ordered curation recipes, where filtering, mixing, and deduplication operators jointly shape the final data distribution. We formulate this problem as fixed-pool data recipe search: given a raw instruction pool and a library of grounded operators, the goal is to discover an executable recipe that constructs a high-quality selected subset under a limited budget of full SFT evaluations, without generating, rewriting, or augmenting training samples. We introduce AutoSelection, a two-layer solver that decouples fixed-pool materialization based on cached task-, data-, and model-side signals from expensive full evaluation, using warmup probes, realized subset states, local recipe edits, Gaussian-process-assisted ranking, and stagnation-triggered reseeding. Experiments on a 90K instruction pool show that AutoSelection achieves the strongest in-distribution reasoning average across three base models, outperforming full-data training, random recipe search, random top-k, and single-operator selectors.

From Instance Selection to Fixed-Pool Data Recipe Search for Supervised Fine-Tuning

Published XAI Agentic AI

EMNLP 2026 · Main Conference

Haodong Wu, Jiahao Zhang, Lijie Hu, Yongqi Zhang

Keywords: Supervised Fine-Tuning, Data Selection, Data Recipe Search, Agentic Optimization, LLM Training

Controlling Repetition in Protein Language Models ICLR 2026 · Poster
Controlling Repetition in Protein Language Models cover Click cover to view abstract
Protein language models (PLMs) frequently collapse into pathological repetition during generation, which undermines structural confidence and functional viability. We present a systematic study of repetition in PLMs and propose Utility-Controlled Contrastive Steering (UCCS), which steers generation using contrastive sets that maximize repetition differences while controlling structural utility. Across ESM-3 and ProtGPT2 on CATH, UniRef50, and SCOP, UCCS reduces repetition without retraining while preserving foldability-related confidence.

Controlling Repetition in Protein Language Models

Published AI4Sci Co-first

ICLR 2026 · Poster

Jiahao Zhang†, Zeqing Zhang†, Di Wang, Lijie Hu

Keywords: Protein Language Models, Reliable Protein Generation, Repetition Control

Bayesian Gated Non-Negative Contrastive Learning ICML 2026
Bayesian Gated Non-Negative Contrastive Learning cover Click cover to view abstract
Accepted to the Forty-Third International Conference on Machine Learning (ICML 2026). A Bayesian gated non-negative contrastive learning framework that improves the interpretability and reliability of learned representations.

Bayesian Gated Non-Negative Contrastive Learning

Published XAI Co-first

ICML 2026

Peng Cui†, Jiahao Zhang†, Lijie Hu

Keywords: Contrastive Learning, Bayesian Gating, Non-Negative Representations, Interpretability

Preprint

Decomposing and Steering Diffusion Transformers with Sparse Autoencoders arXiv · 2026
Decomposing and Steering Diffusion Transformers with Sparse Autoencoders cover Click cover to view abstract
D-Scope connects the interpretation of sparse autoencoder features to generation control in diffusion transformers. It retrieves features from text queries using visual evidence from highly activating image patches, then tests their decoder directions through spatially masked interventions. A study of 150 SAEs across two model families and five layers, together with a benchmark of 100 target concepts, examines reconstruction, feature utilization, visual coverage, and steering effects.

Decomposing and Steering Diffusion Transformers with Sparse Autoencoders

Preprint XAI Co-first

arXiv · 2026

Xinyue Xu†, Jiahao Zhang†, Lijie Hu, Peter Hase, Hao Wang

Keywords: Diffusion Transformers, Sparse Autoencoders, Mechanistic Interpretability, Visual Steering

Under review

Solar-driven evapofiltration enables co-production of lithium and freshwater from extreme brines Submitted to: Nature Sustainability 2026
Solar-driven evapofiltration enables co-production of lithium and freshwater from extreme brines cover Click cover to view abstract
Brines are an important lithium resource, but their utilization remains limited by poor lithium selectivity in highly saline, magnesium-rich environments and by freshwater scarcity in the arid regions where many such resources occur. Here we report evapofiltration, a separation strategy for simultaneous production of freshwater and lithium from extreme brines. The strategy is realized using an evapofiltration membrane that integrates a commercial nanofiltration membrane with a carbon-nanotube photothermal layer in a recirculating NF evaporator. In Dead Sea brine composition, the system driven by solar sustained freshwater production at 1.24 kg/m^2/h while achieving 320-fold lithium enrichment over four stages, ultimately yielding battery-grade Li2CO3 with 99.58% purity. An experimentally informed artificial-intelligence framework further integrates brine chemistry and local solar conditions to predict site-specific lithium and freshwater production from diverse brine resources worldwide. Evapofiltration therefore provides a promising route for integrated lithium recovery and freshwater generation from extreme brines in water-scarce regions.

Solar-driven evapofiltration enables co-production of lithium and freshwater from extreme brines

Under review AI4Sci

Submitted to: Nature Sustainability 2026

Honglang Lu, Jiahao Zhang, Xingxiang Li, Jing Li, Yishuo Huang, Xiaoqin Zhong, Lijie Hu, Jun Ma, Zongyao Zhou

Keywords: Brine Chemistry, Lithium Recovery, Solar Energy, Photothermal Membrane, Evapofiltration, AI for Science

Author mark: † indicates co-first authors.

🚀 Projects

🛠 Open Source Tools

🎖 Honors and Awards

  • 2026.02, MBZUAI Conference Travel Grant.
  • 2024.05, Cloud Computing Application Award, UCB Data Science Discovery Program.
  • 2023 - 2024, Undergraduate Academic Excellence Scholarship.
  • 2023.03, Second Prize, College Student Mathematics Competition (Hubei Division).
  • 2023.01, Second Prize, Chinese Mathematics Competition.
  • 2022.09, Third Prize, China Undergraduate Mathematical Contest in Modeling.

📖 Educations

  • 2025.08 - Present, Ph.D. in Machine Learning, MBZUAI, Abu Dhabi, UAE. Supervisor: Prof. Lijie Hu.
  • 2021.09 - 2025.06, B.Eng. in Electronic Information, Huazhong University of Science and Technology (HUST), Wuhan, China.
  • 2024.01 - 2024.05, Visiting Student, UC Berkeley, Berkeley, CA, USA.

💻 Internships

  • 2025.01 - 2025.07, Research Assistant, Laboratory of Cell Ethology (CIS), Westlake University, Hangzhou, China.
  • 2024.06 - 2024.12, Research Assistant, Representation Learning Lab, Westlake University, Hangzhou, China.
  • 2024.01 - 2024.05, Data Science Research Intern, Grapedata (UC Berkeley), Berkeley, CA, USA.
  • 2023.10 - 2023.12, Research Assistant (Remote), AI Lab (Chaowei Xiao), University of Wisconsin-Madison.

👨‍🏫 Teaching

  • Fall 2026, Teaching Assistant, Generative AI Memorization, MBZUAI.

🧾 Services

  • Conference Reviewer: ICML 2026; NeurIPS 2026; AAAI 2027.
  • Journal Reviewer: IEEE Computational Intelligence Magazine (IEEE CIM); Neurocomputing; ACM Transactions on Probabilistic Machine Learning.