Dynamic Temperature Scheduler for Knowledge Distillation
Sibgat Ul Islam, Jawad Ibn Ahad, Fuad Rahman, Mohammad Ruhul Amin, Nabeel Mohammed, Shafin Rahman
✅ Accepted at IEEE International Conference on Big Data 2025
📄 Paper: arXiv:2511.13767 · IEEE Xplore
Abstract
We introduce the Dynamic Temperature Scheduler (DTS) for knowledge distillation (KD), a method that adaptively adjusts the softening temperature during training based on the cross-entropy loss gap between teacher and student models. Fixed-temperature KD is suboptimal: early training benefits from high T (softer distributions for broad learning) while later training requires low T (sharper distributions for precise mimicry). DTS eliminates manual temperature tuning and integrates seamlessly as a drop-in replacement for any existing KD method. Validated across vision (CIFAR-100, Tiny-ImageNet) and NLP tasks (GLUE, Dolly, SelfInst), DTS consistently improves performance over AKD and CTKD baselines and provides additive gains on top of DKD, MLKD, CRD, and other state-of-the-art KD methods.
Introduction
Knowledge distillation trains a student model to match the teacher’s soft probability distributions. The temperature T controls distribution softness:
- High T → soft distributions → student learns broad inter-class relationships early in training
- Low T → sharp distributions → student learns precise discrimination in later stages
Problem: No single fixed T is optimal across all training steps. Grid search for T is expensive, and heuristic schedules (AKD, CTKD) are not responsive to actual learning dynamics.
DTS Solution: Monitor the teacher-student loss divergence at each batch and adapt T accordingly — no hyperparameter search required.
🎯 Key Contributions
- Loss-Adaptive Temperature — T updated each batch based on divergence
d_loss = L_t − L_sbetween teacher and student cross-entropy losses. - Cosine + Momentum Schedule — Cosine decay provides smooth base annealing; momentum averaging prevents abrupt changes.
- Plug-and-Play — Works with any existing KD method (KD, DKD, MLKD, CRD, FitNet, AT) without architectural changes.
- No Grid Search — Removes costly T hyperparameter sweep entirely.
- Cross-Domain Validation — Vision (CNN, ViT) + NLP (GPT-2, OPT, BERT) tasks validated.
Methodology

Figure: DTS update mechanism. At each training step: (1) compute loss divergence between teacher and student; (2) scale via adaptive α; (3) apply cosine decay; (4) clamp to [T_min, T_max]; (5) apply momentum smoothing for stable temperature update.
DTS Formula
Step 1 — Cosine Scheduling:
p = e_c / e_t (progress ∈ [0,1])
S(p) = 0.5 · (1 + cos(π · p)) (ranges 1.0 → 0.0)
Step 2 — Loss Divergence Adaptive Scaling:
d_loss = L_teacher − L_student
If L_s > L_t (student worse than teacher):
α = d_loss / (d_loss + 1 + β) (β prevents ÷0)
If L_s < L_t (student better):
α = −d_loss / (−d_loss + 1 + β) (still > 0)
Step 3 — Clamped Temperature:
x = T_init · S(p) · α
T_clamped = clamp(x, T_min, T_max)
Step 4 — Momentum Smoothing:
T_current = μ · T_prev + (1 − μ) · T_clamped (μ = 0.9)
📊 Table II — CIFAR-100 Cross-Architecture (Acc %)
| Teacher | Student | AKD | CTKD | KD+DTS |
|---|---|---|---|---|
| VGG13 (74.64) | VGG8 (70.36) | 67.81 | 72.02 | 72.77 |
| ResNet50 (79.34) | MN-V2 (64.60) | 53.86 | 63.17 | 63.24 |
| ResNet56 (72.34) | ResNet20 (69.06) | 68.13 | 69.51 | 70.98 |
| ResNet110 (74.31) | ResNet32 (71.14) | 70.34 | 72.05 | 72.20 |
| VGG13 (74.64) | MN-V2 (64.60) | 52.53 | 62.28 | 62.44 |
| ResNet110 (74.31) | ResNet20 (69.06) | 68.23 | 69.49 | 69.58 |
📊 Table III — CIFAR-100 Same-Architecture, Multiple KD Methods (Acc %)
| Method | ResNet32x4 → ResNet8x4 | ResNet56 → ResNet20 | ResNet110 → ResNet20 | WRN-40-2 → WRN-16-2 | VGG13 → VGG8 | ResNet110 → ResNet32 |
|---|---|---|---|---|---|---|
| Feature Distillation | ||||||
| FitNet | 73.50 | 69.21 | 68.99 | 73.58 | 71.02 | 71.06 |
| AT | 73.44 | 70.55 | 70.65 | 74.08 | 71.43 | 72.31 |
| RKD | 71.90 | 69.61 | 69.25 | 73.35 | 71.48 | 71.82 |
| CRD | 75.51 | 71.16 | 71.46 | 75.48 | 73.94 | 73.48 |
| OFD | 74.95 | 70.98 | 71.29 | 75.24 | 73.95 | 73.23 |
| ReviewKD | 75.63 | 71.89 | 71.34 | 76.12 | 74.84 | 73.89 |
| SimKD | 78.08 | 71.05 | 71.06 | 75.53 | 74.89 | 73.92 |
| CAT-KD | 76.91 | 71.62 | 71.37 | 75.60 | 74.65 | 73.62 |
| Logit Distillation | ||||||
| KD | 72.59 | 69.56 | 69.28 | 73.12 | 70.33 | 71.75 |
| KD + DTS | 72.70 | 70.31 | 69.58 | 73.63 | 72.71 | 72.20 |
| Δ (KD→DTS) | +0.11 | +0.75 | +0.30 | +0.51 | +2.38 | +0.45 |
| DKD | 74.72 | 69.55 | 69.31 | 74.53 | 73.97 | 69.89 |
| DKD + DTS | 74.73 | 69.98 | 69.43 | 74.60 | 74.03 | 70.27 |
| Δ (DKD→DTS) | +0.01 | +0.43 | +0.12 | +0.07 | +0.06 | +0.38 |
| MLKD | 75.03 | 70.85 | 70.32 | 75.18 | 70.52 | 70.32 |
| MLKD + DTS | 75.25 | 70.98 | 70.58 | 75.58 | 71.07 | 70.54 |
| Δ (MLKD→DTS) | +0.22 | +0.13 | +0.26 | +0.40 | +0.55 | +0.22 |
📊 Table IV — CIFAR-100 Cross-Architecture, DTS on Top of KD Methods (Acc %)
| Teacher | Student | Method | Acc % | Δ |
|---|---|---|---|---|
| ResNet32x4 (79.42) | SHN-V2 (71.82) | KD | 71.48 | – |
| KD + DTS | 71.76 | +0.28 | ||
| DKD + DTS | 73.33 | –0.30 | ||
| ResNet50 (79.34) | MN-V2 (64.60) | KD | 62.68 | – |
| KD + DTS | 63.24 | +0.56 | ||
| DKD + DTS | 66.18 | +0.03 | ||
| ResNet32x4 (79.42) | VGG8 (70.36) | KD | 69.95 | – |
| KD + DTS | 71.45 | +1.50 | ||
| MLKD + DTS | 73.35 | +0.62 | ||
| WRN-40-2 (75.61) | SHN-V2 (71.82) | KD | 69.93 | – |
| KD + DTS | 71.35 | +1.42 | ||
| MLKD + DTS | 74.24 | +0.10 | ||
| VGG13 (74.64) | MN-V2 (64.60) | KD | 60.83 | – |
| KD + DTS | 62.44 | +1.61 | ||
| MLKD + DTS | 66.74 | +0.09 |
📊 Table V — Tiny-ImageNet Results (Acc %)
(a) CNN-based Student Models:
| Teacher | Student | KD | KD + DTS | T_max → T_min |
|---|---|---|---|---|
| ResNet34 (65.47) | ResNet18 (60.34) | 57.82 | 60.75 | 4→2 |
| ResNet50 (67.06) | MN-V1 (59.50) | 56.84 | 61.18 | 4→2 |
(b) ViT-based Student Models (Teacher: ResNet50):
| Student | KD | KD + DTS | T_max → T_min |
|---|---|---|---|
| DeiT-Ti | 68.42 | 68.58 | 4→2 |
| T2T-ViT-7 | 64.03 | 64.60 | 4→2 |
| PvT-Ti | 72.01 | 72.69 | 3→1 |
| PiT-Ti | 74.04 | 74.18 | 3→1 |
📊 Table VI — NLP Tasks: ROUGE-L Scores
| Teacher | Student (Params) | Method | Dolly | SelfInst | S-NI | UnNI | Vicuna |
|---|---|---|---|---|---|---|---|
| GPT-2 (1.5B) | — | Teacher | 27.6 | 14.3 | 27.6 | 31.8 | 16.3 |
| GPT-2 (120M) | SFT w/o KD | 23.3 | 10.0 | 14.7 | 18.5 | 16.3 | |
| KD | 21.88 | 9.74 | 14.88 | 18.46 | 13.92 | ||
| KD + DTS | 21.48 | 9.82 | 16.16 | 19.66 | 13.95 | ||
| SeqKD | 22.27 | 9.47 | 14.83 | 17.37 | 14.90 | ||
| SeqKD + DTS | 21.30 | 10.89 | 15.54 | 18.77 | 14.15 | ||
| OPT (1.3B) | — | Teacher | 26.0 | 11.4 | 23.1 | 28.4 | 15.6 |
| OPT (125M) | KD | 19.41 | 8.06 | 15.30 | 17.52 | 13.65 | |
| KD + DTS | 19.91 | 8.46 | 15.64 | 19.66 | 14.05 | ||
| SeqKD | 20.14 | 8.24 | 15.52 | 17.98 | 14.64 | ||
| SeqKD + DTS | 20.33 | 8.95 | 16.74 | 18.25 | 13.58 |
📊 Table VII — GLUE Tasks: TinyBERT-4L-312D from BERT_base_uncased
| Method | CoLA (MCC) | MRPC (Acc) | RTE (Acc) |
|---|---|---|---|
| SFT | 0.52 | 87.00 | 72.00 |
| TinyBERT | 0.35 | 83.57 | 67.14 |
| with OP | 0.36 | 84.31 | 67.87 |
| OE + OP | 0.39 | 84.33 | 66.06 |
📚 Citation
@inproceedings{islam2025dts,
title={Dynamic Temperature Scheduler for Knowledge Distillation},
author={Islam, Sibgat Ul and Ahad, Jawad Ibn and Rahman, Fuad and Amin, Mohammad Ruhul and Mohammed, Nabeel and Rahman, Shafin},
booktitle={IEEE International Conference on Big Data (BigData)},
year={2025}
}
