CANLI
xAI, Imagine API’yi 2.0’a Yükseltmeye Hazırlanıyor: Görüntü ve Video Tek…·Microsoft MAI-Cyber-1-Flash’ı Duyurdu·Moonshot AI, Kimi K3 Model Ağırlıklarını ve Teknik Raporunu Açık…
29 Sep 2026 · 15:43 GMT+3
Ai Haber – Türkiyenin Yapay Zeka Haber Portalı
ARAşTıRMA · MAKINE ÖğRENMESI arXiv:2609.35763 28 Eyl 2026 · v1

Unifying Distributional Training for One-Step Visual Generation

Chi Zhang, Haoyang Shi, Yueyi Liu, Ruichuan An, Junkang Zhou, +7 yazar

YAYIN:28 Eyl 2026 ALAN:cs.LG OKUMA:2

Özet

emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates textbf{MGFlow}, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with textbf{1.45} $mathrm{FDr}^6$ on pMF-H and textbf{1.64} on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. Project page: https://shihaoyang0423.github.io/MGFlow-website/

Özetle: emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces.

Özet

emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates textbf{MGFlow}, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with textbf{1.45} $mathrm{FDr}^6$ on pMF-H and textbf{1.64} on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. Project page: https://shihaoyang0423.github.io/MGFlow-website/

Orijinal Özet (İngilizce)

emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates textbf{MGFlow}, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with textbf{1.45} $mathrm{FDr}^6$ on pMF-H and textbf{1.64} on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. Project page: https://shihaoyang0423.github.io/MGFlow-website/

Kaynak: arXiv:2609.35763 · PDF

BibTeX

@article{zhang2026unifying,
  title   = {Unifying Distributional Training for One-Step Visual Generation},
  author  = {Chi Zhang and Haoyang Shi and Yueyi Liu and Ruichuan An and Junkang Zhou and Chang Li and Xiuyuan Lu and Yichi Zhang and Bo Wang and Yuhang Wu and Sen Cui and Miao Liu},
  journal = {arXiv preprint arXiv:2609.35763},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.35763}
}

Tartışma

Bu habere emoji ile tepki ver

Hizli:

Henüz yorum yok. İlk yorumu siz yapın!

Yapıcı ve saygılı yorumlar bekliyoruz. Topluluk kuralları