The project demonstrates a markerless motion-capture and gesture-recognition pipeline for mobile and desktop applications. A live camera feed is processed frame by frame to extract human body landmarks, draw a tracking skeleton, and recognize actions including walking, clapping, squatting, kicking, hand gestures, and idle movement.
The resulting joint coordinates and action states are transmitted into Unity, where a C# retargeting layer maps the performer’s movement onto a rigged Japanese anime character. The architecture can support interactive VTuber avatars, fitness applications, games, virtual production, and markerless AR or VR controls without specialist mocap hardware.
Benefits
Capture full-body movement without suits, markers, or specialist depth hardware
Enable interactive avatars and VTuber experiences from a standard camera
Reuse pose and gesture signals for fitness, games, AR, and VR interaction
Support real-time Unity experiences through a modular vision-to-engine pipeline
Features
Frame-by-frame MediaPipe Pose landmark extraction and skeleton overlay
Real-time classification of walking, clapping, squatting, kicking, gestures, and idle states
3D joint-coordinate processing through Python and OpenCV
Unity C# interoperability with live bone-transform retargeting
Markerless tracking architecture for conventional mobile and desktop cameras
Real-Time AI Motion Capture in Unity 3D Responsive, markerless full-body motion and gesture retargeting from a standard camera to a Unity 3D avatar
Year
2026
Industry
Interactive Entertainment, Fitness & Virtual Production
Outcome
Responsive, markerless full-body motion and gesture retargeting from a standard camera to a Unity 3D avatar
Quick answer
Need a Computer Vision developer?
Fareed Tahir designs and develops production-focused immersive experiences in Unity, AR, VR, WebAR and interactive 3D. Remote collaboration is available worldwide from Lahore, Pakistan.