What Does Privacy Actually Cost? Tuning Differentially Private Image Classifiers
If you train models on data about people, the model itself can leak the data it was trained on.
A trained network is not a black box in the privacy sense. Given access to a model, an adversary can often infer whether a particular record was in the training set, and sometimes reconstruct fragments of it. That matters most in exactly the domains where machine learning is most useful: medical imaging, biometrics, and any dataset where records belong to real people.
Differential privacy (DP) is the standard answer. It guarantees that the output of a computation is nearly unchanged whether or not any single individual’s data was included. The guarantee is not free.
The Privacy-Accuracy Tradeoff
Making deep learning differentially private means injecting calibrated noise into training itself. The standard technique, DP-SGD, clips each per-example gradient to a fixed norm, then adds Gaussian noise before the update.
Both steps hurt learning. Clipping discards gradient magnitude; noise obscures the signal the model is trying to follow. The privacy budget epsilon (ε) controls how much noise is required; a smaller ε means a stronger guarantee and a less accurate model.
That framing is correct but too coarse to be useful. In practice, ε is only one of several parameters that decide whether a private model is usable, and two teams training under the same ε can land tens of accuracy points apart. So the real question is:
Given a fixed privacy budget, which training choices determine how much accuracy you keep?
The Experiments
This started as my final project for CSCI 8960: Privacy-Preserving Data Analysis at the University of Georgia. The setup was deliberately broad rather than deep — many models under many privacy settings, rather than one architecture tuned hard.
- Models: ConvNet, ResNet18, EfficientNet, DenseNet121, and ViT, plus SVM, KNN, and Naive Bayes as classical contrasts
- Data: CIFAR-10 under (ε, δ)-DP, with δ fixed at 1e-5
- Tooling: PyTorch and Opacus for DP-SGD, on Colab and Kaggle GPUs
- Scope: 70+ runs sweeping ε, clipping threshold, optimizer, batch size, learning rate, and epochs
What Mattered Most
-
Privacy budget dominates. Tightening ε from 20 to 1 degraded accuracy for every architecture. This is the trade-off in its clearest form, and it anchors everything else. EfficientNet held the top spot at every budget level.
-
Clipping threshold is a real lever Raising it from 1 to 5 improved accuracy across models, since more of the gradient magnitude survives before noise is added. Stated carefully: this held within the range tested. Noise scales with the threshold, so the relationship is not monotonic in general; clipping is a parameter to tune, not a constant to accept.
-
Optimizer choice matters more under DP than without it. Adam consistently beat SGD, RMSProp, and Adagrad. DP-SGD produces gradients that are both noisy and clipped, which is exactly the regime where adaptive per-parameter step sizes help.
-
Larger batches help. Going from 128 to 256 improved accuracy everywhere. Noise is added once per batch, so bigger batches improve the signal-to-noise ratio of each update.
-
Learning rate is easy to get wrong. 1e-2 degraded performance badly; 1e-3 and 5e-4 performed comparably. Ordinary advice, but the penalty is sharper when gradients are already perturbed.
-
More epochs helped, up to a point. Training longer usually improved accuracy, but under a fixed budget, more epochs mean more noisy gradient releases and a larger noise multiplier per step. Extra training is not free.
The classical classifiers performed poorly, but that is not a clean head-to-head: they were trained on Laplace-perturbed inputs rather than with DP-SGD. The defensible conclusion is that input perturbation is a poor privacy mechanism for image classification.
Where This Landed
The best configuration was EfficientNet at 59.63% test accuracy on CIFAR-10 under ε = 5.0 (Adam, batch size 256, 100 epochs, clipping threshold 1.0, noise multiplier 0.912).
Against non-private CIFAR-10 baselines that comfortably exceed 90%, that number makes the cost of privacy concrete. It is also below what later work has achieved with better pretraining and larger effective batches.
The more durable result is the shape of the trade-off, not the number. Privacy cost is not a fixed tax set by ε alone; a meaningful share of the accuracy lost under DP is recoverable through choices that have nothing to do with the guarantee itself. Teams that treat DP as a switch to flip tend to conclude private training does not work. Teams that treat it as a regime with its own tuning get much further.
The limits are real: compute was course-scale, configurations were repeated at most three times, and the largest models could not be run at every batch size. This is an empirical map of the parameter space, not a theoretical result. But the map is the useful part.
Learn More
Implementation, training scripts and experiment configs are available on GitHub. Experimental results and other details can be found in this ArXiv pre-print.
This work was completed as a course project for CSCI 8960: Privacy-Preserving Data Analysis, University of Georgia.