【主讲人简介】:魏鸿鑫,南方科技大学pc蛋蛋系助理教授,主要研究机器学习不确定性量化,及其在数据优化与隐私中的应用。他的研究致力于使机器学习模型能够准确表达预测中的不确定性,为可信推断与高效训练提供原则指导。魏鸿鑫老师分别于华中科技大学和新加坡南洋理工大学获得学士和博士学位,曾在清华交叉信息研究院担任研究助理并在美国威斯康辛大学麦迪逊分校进行研究访问。其获得了深圳市面上、广东省面上等项目资助并以项目骨干参与了国家重点研发计划“数学和应用研究”重点专项青年科学家项目。他近年已在ICML, NeurIPS, ICLR等国际顶级会议和JMLR等期刊发表论文 57 篇。其受邀担任 ICML、NeurIPS、ICLR 等国际机器学习会议领域主席,以及 JASA、JMLR、TPAMI、IJCV 等顶级期刊审稿人。在开源项目方面,其团队开发并开源了深度学习共形预测工具库TorchCP (社区下载安装量逾 24000 次)和大模型中文内容安全评测基准 ChineseSafe(下载量逾 32000 次)。
【内容简介】:Identifying training data of large-scale models is critical for copyright litigation, privacy auditing, and ensuring fair evaluation. However, existing works typically treat this task as an instance-wise identification without controlling the error rate of the identified set, which cannot provide statistically reliable evidence. In this work, we formalize training data identification as a set-level inference problem and propose Provable Training Data Identification (PTDI), a distribution-free approach that enables provable and strict false identification rate control. Specifically, our method computes conformal p-values for each data point using a set of known unseen data and then develops a novel Jackknife-corrected Beta boundary (JKBB) estimator to estimate the training-data proportion of the test set, which allows us to scale these p-values. By applying the Benjamini–Hochberg (BH) procedure to the scaled p-values, we select a subset of data points with provable and strict false identification control. Extensive experiments across various models and datasets demonstrate that PTDI achieves higher power than prior methods while strictly controlling the FIR.
【讲座时间】:2026年8月11日(星期二)下午14:00
【讲座地点】:人文社科科研楼1801会议室



会议室预约
资料下载