单目 RGB 手部姿态估计新框架 CS-ViT:HO3D 上超 SOTA 37.1%
Estimating Accurate Hand Pose in Camera Space with Vision Transformer
单目 RGB 就能算 3D 手部姿态,HO3D 上比 SOTA 准最多 37.1%,做手势交互的可以看代码。
论文提出 CS-ViT 框架,解决单目 RGB 手部姿态估计中的深度歧义,以及局部手姿与腕部位置在透视投影下的耦合问题。框架用变换同构监督提取手部深度信息,用透视信息嵌入解耦局部姿态与腕部位置,整体建立在主流 encoder-decoder 架构上。训练侧引入帧率感知的多数据集策略做序列姿态优化。在 HO3D 基准上,CS-MJE 指标最高较 SOTA 领先 37.1%,代码开源于 GitHub 的 Mine268/CS-ViT 仓库。
Estimating Accurate Hand Pose in Camera Space with Vision Transformer
Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. However, this camera-space estimation confronts two fundamental challenges: (1) depth ambiguity in monocular settings, and (2) the coupling effect of hand local poses and global wrist positions in the perspective projections. In particular, this coupling reflects that the projections are jointly determined by local hand poses, wrist positions, and camera intrinsics. To overcome these challenges, our framework proposes two key innovations: Transformation-Isomorphism Supervision for hand-depth information extraction and Perspective Information Embedding for resolving above coupling effect of local pose and wrist position, both integrated within the mainstream encoder-decoder architecture. Besides, we propose a novel framerate-aware multi-dataset training strategy for sequential pose refinement. Our fully integrated approach achieves at most 37.1\% superiority in CS-MJE over SOTA on HO3D. Project page: https://github.com/Mine268/CS-ViT.