
Researchers at the Max Planck Institute for Intelligent Systems and Carnegie Mellon University have developed MAMMA, a motion-capture system that reconstructs human movement from multicamera video without physical tracking markers. Its advantage is handling close interactions, including hugs, dancing, and martial arts, where people obscure one another from view, tells Tech Xplore.
MAMMA, short for Markerless Accurate Multi-person Motion Acquisition, estimates 512 virtual points on the body from video frames. It matches these points across camera views and fits the SMPL-X digital human model to recover body shape, posture, and hand movement. The underlying neural network also estimates visibility, uncertainty, and contact, helping distinguish interacting people and reconstruct plausible motion.
Traditional motion capture requires specialized studios, marker-covered suits, and manual cleanup. Markers can fall off or become hidden during physical contact, making reconstruction difficult. Commercial markerless alternatives also tend to be expensive and can struggle with complex movements.
The researchers report that MAMMA’s pose accuracy differed from a conventional suit-and-marker system by less than one millimeter in their comparison. This describes the difference between systems, rather than an overall reconstruction error below one millimeter. They also say processing was nearly three times faster than the traditional manual workflow. Demonstrations used consumer equipment, including multiple iPhones.
Potential applications extend beyond film and games. Physical therapists could monitor patients without intrusive markers, while biomechanics researchers could study movement in natural settings. Coaches could analyze several athletes without bulky tracking gear. Virtual reality developers could create expressive avatars with more realistic body and hand interactions.
For smaller studios, capturing wrestling, dancing, and other contact-heavy animation without a dedicated facility could lower production barriers. The team is sharing research data and tools with the academic community to encourage further development. Presented at CVPR 2026, the work shows how computer vision can make detailed motion capture more accessible.
