CANLI
xAI, Imagine API’yi 2.0’a Yükseltmeye Hazırlanıyor: Görüntü ve Video Tek…·Microsoft MAI-Cyber-1-Flash’ı Duyurdu·Moonshot AI, Kimi K3 Model Ağırlıklarını ve Teknik Raporunu Açık…
8 Oct 2026 · 23:16 GMT+3
Ai Haber – Türkiyenin Yapay Zeka Haber Portalı
ARAşTıRMA · BILGISAYARLı GöRü arXiv:2610.10538 7 Eki 2026 · v1

Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos

Shravan Chaudhari, William Paul, Suchi Saria, Rama Chellappa, Homanga Bharadhwaj

YAYIN:7 Eki 2026 ALAN:cs.CV OKUMA:6

Özet

As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person's day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object's observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise. Short descriptions preserve details such as an object's contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementary roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.

Özetle: As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later.

Özet

As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person's day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object's observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise. Short descriptions preserve details such as an object's contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementary roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.

Orijinal Özet (İngilizce)

As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person's day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object's observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise. Short descriptions preserve details such as an object's contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementary roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.

Kaynak: arXiv:2610.10538 · PDF

BibTeX

@article{chaudhari2026never,
  title   = {Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos},
  author  = {Shravan Chaudhari and William Paul and Suchi Saria and Rama Chellappa and Homanga Bharadhwaj},
  journal = {arXiv preprint arXiv:2610.10538},
  year    = {2026},
  url     = {https://arxiv.org/abs/2610.10538}
}

Tartışma

Bu habere emoji ile tepki ver

Hizli:

Henüz yorum yok. İlk yorumu siz yapın!

Yapıcı ve saygılı yorumlar bekliyoruz. Topluluk kuralları