Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation

Liu, Huan; Chen, Qiang; Tan, Zichang; Liu, Jiang-Jiang; Wang, Jian; Su, Xiangbo; Li, Xiaolong; Yao, Kun; Han, Junyu; Ding, Errui; Zhao, Yao; Wang, Jingdong

Abstract:In this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.g., regarding pose estimation as keypoint box detection and combining with human detection in ED-Pose, hierarchically predicting with pose decoder and joint (keypoint) decoder in PETR. We present a simple yet effective transformer approach, named Group Pose. We simply regard $K$-keypoint pose estimation as predicting a set of $N\times K$ keypoint positions, each from a keypoint query, as well as representing each pose with an instance query for scoring $N$ pose predictions. Motivated by the intuition that the interaction, among across-instance queries of different types, is not directly helpful, we make a simple modification to decoder self-attention. We replace single self-attention over all the $N\times(K+1)$ queries with two subsequent group self-attentions: (i) $N$ within-instance self-attention, with each over $K$ keypoint queries and one instance query, and (ii) $(K+1)$ same-type across-instance self-attention, each over $N$ queries of the same type. The resulting decoder removes the interaction among across-instance type-different queries, easing the optimization and thus improving the performance. Experimental results on MS COCO and CrowdPose show that our approach without human box supervision is superior to previous methods with complex decoders, and even is slightly better than ED-Pose that uses human box supervision. $\href{this https URL}{\rm Paddle}$ and $\href{this https URL}{\rm PyTorch}$ code are available.

Comments:	Accepted by ICCV 2023
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2308.07313 [cs.CV]
	(or arXiv:2308.07313v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2308.07313

Computer Science > Computer Vision and Pattern Recognition

Title:Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators