Provable Trustworthy Machine Learning through Differential Trust

Trustworthy machine learning faces several fundamental challenges. First, many trust notions are semantically difficult to define: memorization, data poisoning, and copyright protection depend on context, threat models, and downstream interpretation, rather than a single universal metric. Second, quantifying the risk for modern scaling, black-box machine learning models is difficult. Third, operationally, it largely remains open how to transform an ordinary training procedure into one with provable trust guarantees while preserving model utility.
In this talk, I will present Differential Trust, a unified framework that addresses these challenges through reference indistinguishability. Instead of explicitly characterizing every possible undesirable behavior, the key idea is to identify safe reference models — such as models trained without a sensitive, malicious, or protected subset of data — and ensure that the target model is indistinguishable from these references.
I will discuss two approaches within this framework. The first, Data-Specific Indistinguishability (DaSI), is a noise-based approach that calibrates randomization to the specific target-reference differences, providing sharper utility than worst-case indistinguishability. The second, Distribution-Specific Indistinguishability (DiSI), shows that when the population data distribution provides sufficient randomness, one can even remove additional noise perturbation and obtain trust guarantees exploiting input randomness in an adversarial setting.
Together, these works provide a systematic approach to memorization mitigation, backdoor defense, and copyright protection through provable control of data influence.
Hanshen Xiao is an Assistant Professor of Computer Science at Purdue University and a Faculty Researcher at NVIDIA. He received his Ph.D. in Computer Science from MIT and his B.S. in Mathematics from Tsinghua University.
His research focuses on provable trustworthy machine learning and computation, including automated black-box privatization, differential trust, and adversarial robustness. He is a recipient of the MathWorks Fellowship and the Tsinghua Future Scholar Fellowship.