Process / pipelineBiomechanicsComputer visionPipeline

Markerless Motion Capture

Also known as: Marker-free tracking, Vision-based motion capture, Deep learning pose estimation

OriginatorZhe CaoYear2017Sources2Related methods5

Markerless motion capture infers the 3D positions and joint angles of a moving subject from video sequences using computer vision and machine learning. Pioneered by deep learning approaches such as OpenPose and MediaPipe, it eliminates the need for reflective markers or inertial sensors, making motion capture accessible and practical for real-world applications.

Key highlights

  • No specialized hardware or markers required; uses standard cameras
  • Practical for real-time applications and field deployments
  • Handles multiple people in a single view
  • Improving accuracy as deep learning models mature and datasets expand

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use markerless motion capture when markers are impractical (large crowds, outdoor settings, water immersion) or when rapid deployment is required. It excels in fitness, sports analysis, clinical screening, and entertainment. Assumptions include sufficient image resolution to resolve joints, adequate lighting, and pre-training on diverse poses. Be cautious with heavy occlusion, fast motion blur, or extreme poses outside the training distribution.

Strengths & limitations

Strengths
  • No specialized hardware or markers required; uses standard cameras
  • Practical for real-time applications and field deployments
  • Handles multiple people in a single view
  • Improving accuracy as deep learning models mature and datasets expand
Limitations
  • Accuracy degrades with occlusion, motion blur, and poor lighting
  • Lower spatial resolution than marker-based systems in clinical settings
  • Temporal noise and jitter without post-processing or smoothing
  • Struggles with extreme poses or subjects outside the training distribution

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How accurate is markerless motion capture compared to marker-based systems?

In controlled laboratory settings with good lighting, state-of-the-art markerless systems achieve 2–5 cm 3D error at major joints; marker-based systems achieve sub-centimeter accuracy. Markerless degrades more in unconstrained environments.

Can I improve accuracy with multi-view markerless capture?

Yes. Multiple synchronized cameras allow 3D triangulation and reduce ambiguity from self-occlusion. Accuracy improves especially when combining 2D detections from different viewpoints.

How do I reduce temporal jitter in markerless tracking?

Apply temporal smoothing (Kalman filtering, butterworth filtering, or optical flow) post-hoc to the joint trajectories. Some frameworks include smoothing as a built-in option.

Sources

  1. 1.
    Cao, Z., Simon, T., Wei, S. E., & Sheikh, Y. (2017). Realtime multi-person 2D pose estimation using part affinity fields. IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  2. 2.
    Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Markerless Motion Capture. ScholarGate. https://scholargate.app/biomechanics/markerless-motion-capture