In-Distribution Public Data Synthesis With Diffusion Models for Differentially Private Image Classification

Jinseong Park, Yujin Choi, Jaewook Lee
2/9/2026

Abstract

To alleviate the utility degradation of deep learning image classification with differential privacy (DP), employing extra public data or pre-trained models has been widely explored. Recently, the use of in-distribution public data has been investigated, where tiny subsets of datasets are released publicly. In this paper, we investigate a framework that leverages recent diffusion models to amplify the information of public data. Subsequently, we identify data diversity and generalization gap between public and private data as critical factors addressing the limited public data. While assuming 4% of training data as public, our method achieves 85.48% on CIFAR-10 with a privacy budget of $\varepsilon, =2$, without employing extra public data for training.

DOISemantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 3 weaker matches are hidden by default — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{park2026indistribution,
  title  = {In-Distribution Public Data Synthesis With Diffusion Models for Differentially Private Image Classification},
  author = {Jinseong Park and Yujin Choi and Jaewook Lee},
  year   = {2026},
  doi    = {10.1109/CVPR52733.2024.01163},
  url    = {https://doi.org/10.1109/CVPR52733.2024.01163},
  journal = {CVPR 2024 2024}
}

Discussion