ACOI: Reconstructing Adult–Child–Object Interaction from in-the-wild video

University at Buffalo

Abstract

Understanding how adults and young children play together requires more than tracking one person at a time. Current 3D human models are trained mostly on adults and handle one person at a time. When applied to real home videos, they turn children into shrunken adults, produce jittery hands, and let bodies pass through each other and through the objects they hold. This project builds a system prototype that reconstructs, from a single monocular video, the adult, the child, the furniture, and the hand-held toy in one shared 3D space. It aims for an age-accurate child, precise hand gestures, correct hand–object contact, and no interpenetration.