CANLI
xAI, Imagine API’yi 2.0’a Yükseltmeye Hazırlanıyor: Görüntü ve Video Tek…·Microsoft MAI-Cyber-1-Flash’ı Duyurdu·Moonshot AI, Kimi K3 Model Ağırlıklarını ve Teknik Raporunu Açık…
8 Oct 2026 · 22:29 GMT+3
Ai Haber – Türkiyenin Yapay Zeka Haber Portalı
ARAşTıRMA · BILGISAYARLı GöRü arXiv:2610.10512 7 Eki 2026 · v1

Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery

Chen Xu, Yunqi Li, Binbin Huang, Brent Yi, Shenghua Gao, +1 yazar

YAYIN:7 Eki 2026 ALAN:cs.CV OKUMA:10

Özet

Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent. To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions. Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence. The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency. Their learned 2D-3D correspondence further enables recovery of the hand's global position and orientation relative to the camera. Extensive experiments on challenging benchmarks demonstrate significantly improved accuracy and speed in local hand-pose and camera-space reconstruction. Notably, our method captures much better hand-motion dynamics, producing significantly smoother motion than previous methods while maintaining high per-frame pose accuracy.

Özetle: Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent.

Özet

Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent. To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions. Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence. The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency. Their learned 2D-3D correspondence further enables recovery of the hand's global position and orientation relative to the camera. Extensive experiments on challenging benchmarks demonstrate significantly improved accuracy and speed in local hand-pose and camera-space reconstruction. Notably, our method captures much better hand-motion dynamics, producing significantly smoother motion than previous methods while maintaining high per-frame pose accuracy.

Orijinal Özet (İngilizce)

Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent. To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions. Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence. The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency. Their learned 2D-3D correspondence further enables recovery of the hand's global position and orientation relative to the camera. Extensive experiments on challenging benchmarks demonstrate significantly improved accuracy and speed in local hand-pose and camera-space reconstruction. Notably, our method captures much better hand-motion dynamics, producing significantly smoother motion than previous methods while maintaining high per-frame pose accuracy.

Kaynak: arXiv:2610.10512 · PDF

BibTeX

@article{xu2026videoconditioned,
  title   = {Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery},
  author  = {Chen Xu and Yunqi Li and Binbin Huang and Brent Yi and Shenghua Gao and Yi Ma},
  journal = {arXiv preprint arXiv:2610.10512},
  year    = {2026},
  url     = {https://arxiv.org/abs/2610.10512}
}

Tartışma

Bu habere emoji ile tepki ver

Hizli:

Henüz yorum yok. İlk yorumu siz yapın!

Yapıcı ve saygılı yorumlar bekliyoruz. Topluluk kuralları