Research on a Transformer Algorithm with Mixture-of-Experts for 3D Human Pose Estimation
DOI:
https://doi.org/10.54097/fp1ba208Keywords:
3D Human Pose Estimation, Transformer, Mixture of Experts, Dynamic Routing, Shared ExpertsAbstract
Monocular 3D Human Pose Estimation Technology has been widely applied in the field of intelligent rehabilitation devices. However, existing Transformer-based methods typically adopt a uniform parameter structure, making it difficult to effectively capture the significant pose feature differences among various action categories. This leads to suboptimal performance when handling complex or unseen actions. In recent years, the Mixture-of-Experts (MoE) architecture has achieved breakthrough progress in large-scale models, offering new insights for human pose estimation. Inspired by DeepSeekMoE, this paper proposes a novel 3D human pose estimation framework, PoseMoE Former, which significantly enhances the model’s generalization ability and modeling efficiency across different action categories by integrating both shared expert and routing expert mechanisms. Specifically, our model comprises three core modules: static and dynamic routing mechanisms that enable expert selection; shared expert and routing expert modules that capture general low-level features and category-specific features; and a loss system including the main task loss, dynamic loss, and load balancing loss, which alleviates issues of unbalanced expert utilization. Experiments on Human3.6M demonstrate that PoseMoE Former outperforms mainstream baseline methods, particularly excelling in complex action categories. Ablation studies further confirm the rationality of each module’s design, showing notable academic value and practical application potential.
Downloads
References
[1] Huang, Q., An, Z., & Zhuang, N., et al. (2024). Harder tasks need more experts. arXiv preprint arXiv:2403.07652.
[2] Hossain, M. R. I., & Little, J. J. (2018). Exploiting temporal information for 3D human pose estimation. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 68–84).
[3] Zuo, S., Zhang, Q., Liang, C., et al. (2022). Moebert: From BERT to mixture-of-experts via importance-guided adaptation. arXiv preprint arXiv:2204.07675.
[4] Fedus, W., Zoph, B., & Shazeer, N. (2022). Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(1), 1–39. https://jmlr.org/papers/v23/21-1399.html
[5] Peng, J., Zhou, Y., & Mok, P. Y. (2024). KTPFormer: Kinematics and trajectory prior knowledge enhanced transformer for 3D human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
[6] Mehta, D., Rhodin, H., Casas, D., et al. (2017). Monocular 3D human pose estimation in the wild using improved CNN supervision. In Proceedings of the International Conference on 3D Vision (3DV) (pp. 506–516). IEEE. https://ieeexplore.ieee.org/document/8283659. https://doi.org/10.1109/3DV.2017.00064
[7] Mehta, D., Rhodin, H., Casas, D., et al. (2017). Monocular 3D human pose estimation in the wild using improved CNN supervision. In Proceedings of the International Conference on 3D Vision (3DV) (pp. 506–516). IEEE. https://ieeexplore.ieee.org/document/8283659. https://doi.org/10.1109/3DV.2017.00064
[8] Martinez, J., Hossain, R., Romero, J., et al. (2017). A simple yet effective baseline for 3D human pose estimation. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) (pp. 2640–2649).
[9] Martinez, J., Hossain, R., Romero, J., et al. (2017). A simple yet effective baseline for 3D human pose estimation. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) (pp. 2640–2649).
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Computer Science and Artificial Intelligence

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.








