CANLI
xAI, Imagine API’yi 2.0’a Yükseltmeye Hazırlanıyor: Görüntü ve Video Tek…·Microsoft MAI-Cyber-1-Flash’ı Duyurdu·Moonshot AI, Kimi K3 Model Ağırlıklarını ve Teknik Raporunu Açık…
9 Oct 2026 · 00:11 GMT+3
Ai Haber – Türkiyenin Yapay Zeka Haber Portalı
ARAşTıRMA · İSTATISTIKSEL ÖğRENME arXiv:2610.10488 7 Eki 2026 · v1

Best Arm Identification for Bandits with Shifting Means

Lukas Zierahn, Wouter M. Koolen, Shubhada Agrawal, Christina Katsimerou, Dirk van der Hoeven

YAYIN:7 Eki 2026 ALAN:stat.ML

Özet

We study the best arm identification problem in a stochastic environment with a novel form of adversarial perturbations, which we coin Shifting Means. While classically the mean rewards of the $K$ arms are stable in time, in Shifting Means only the gaps $boldsymbolΔ$ between mean rewards are stable, while their common shift may be determined adversarially in each round. The objective of the learner is to identify the best arm with high probability while minimizing sample complexity (the fixed confidence setting). Handling shifts requires new tools: we show that algorithms employing a Generalized Likelihood Ratio Test (GLRT) stopping rule, including the popular Track-and-Stop, fail under time-varying shifts. Instead, we propose Importance Weights for Shifting Means ($mathsf{ISM}$). Assuming means bounded by $U$ and $σ^2$-sub-Gaussian rewards, we show $mathsf{ISM}$ to be $δ$-correct and to enjoy a sample complexity bound of order $K (σ^2 + U^2) Δ_{min}^{-2} ln frac{1}δ$. We also present a matching (up to constant factors) worst-case lower bound and evaluate our results empirically.

Özetle: We study the best arm identification problem in a stochastic environment with a novel form of adversarial perturbations, which we coin Shifting Means.

Özet

We study the best arm identification problem in a stochastic environment with a novel form of adversarial perturbations, which we coin Shifting Means. While classically the mean rewards of the $K$ arms are stable in time, in Shifting Means only the gaps $boldsymbolΔ$ between mean rewards are stable, while their common shift may be determined adversarially in each round. The objective of the learner is to identify the best arm with high probability while minimizing sample complexity (the fixed confidence setting). Handling shifts requires new tools: we show that algorithms employing a Generalized Likelihood Ratio Test (GLRT) stopping rule, including the popular Track-and-Stop, fail under time-varying shifts. Instead, we propose Importance Weights for Shifting Means ($mathsf{ISM}$). Assuming means bounded by $U$ and $σ^2$-sub-Gaussian rewards, we show $mathsf{ISM}$ to be $δ$-correct and to enjoy a sample complexity bound of order $K (σ^2 + U^2) Δ_{min}^{-2} ln frac{1}δ$. We also present a matching (up to constant factors) worst-case lower bound and evaluate our results empirically.

Orijinal Özet (İngilizce)

We study the best arm identification problem in a stochastic environment with a novel form of adversarial perturbations, which we coin Shifting Means. While classically the mean rewards of the $K$ arms are stable in time, in Shifting Means only the gaps $boldsymbolΔ$ between mean rewards are stable, while their common shift may be determined adversarially in each round. The objective of the learner is to identify the best arm with high probability while minimizing sample complexity (the fixed confidence setting). Handling shifts requires new tools: we show that algorithms employing a Generalized Likelihood Ratio Test (GLRT) stopping rule, including the popular Track-and-Stop, fail under time-varying shifts. Instead, we propose Importance Weights for Shifting Means ($mathsf{ISM}$). Assuming means bounded by $U$ and $σ^2$-sub-Gaussian rewards, we show $mathsf{ISM}$ to be $δ$-correct and to enjoy a sample complexity bound of order $K (σ^2 + U^2) Δ_{min}^{-2} ln frac{1}δ$. We also present a matching (up to constant factors) worst-case lower bound and evaluate our results empirically.

Kaynak: arXiv:2610.10488 · PDF

BibTeX

@article{zierahn2026best,
  title   = {Best Arm Identification for Bandits with Shifting Means},
  author  = {Lukas Zierahn and Wouter M. Koolen and Shubhada Agrawal and Christina Katsimerou and Dirk van der Hoeven},
  journal = {arXiv preprint arXiv:2610.10488},
  year    = {2026},
  url     = {https://arxiv.org/abs/2610.10488}
}

Tartışma

Bu habere emoji ile tepki ver

Hizli:

Henüz yorum yok. İlk yorumu siz yapın!

Yapıcı ve saygılı yorumlar bekliyoruz. Topluluk kuralları