Single-View 6D Object Pose Estimation with Deep Learning
Cai, Dingding (2026)
Cai, Dingding
Tampere University
2026
Tieto- ja sähkötekniikan tohtoriohjelma - Doctoral Programme in Computing and Electrical Engineering
Informaatioteknologian ja viestinnän tiedekunta - Faculty of Information Technology and Communication Sciences
This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited.
Väitöspäivä
2026-08-28
Julkaisun pysyvä osoite on
https://urn.fi/URN:ISBN:978-952-03-4697-3
https://urn.fi/URN:ISBN:978-952-03-4697-3
Tiivistelmä
Single-view six-degree-of-freedom (6D) object pose estimation, which involves determining the 3D orientation and position of an object from a single RGB or RGB-D (RGB plus Depth) image, remains a longstanding research challenge in computer vision. Accurate 6D pose estimation is crucial for various applications, including augmented reality, robotic manipulation, and autonomous driving. Despite significant advancements in deep learning-based 6D object pose estimation, several challenges persist: (1) pose ambiguity, where multiple distinct poses correspond to visually equivalent observations; (2) domain shift, where the data distribution of the testing phase significantly differs from that of training; and (3) model generalizability, the ability to handle open-set objects not seen during training. This dissertation addresses these challenges by exploring four key research questions and contributing novel deep learning-based solutions.
Firstly, this dissertation introduces a deep neural network model (SC6D) designed to address pose ambiguity. SC6D maps RGB images and object poses into a shared latent space, enabling the implicit handling of ambiguous poses through learned representations. Building on SC6D, a monocular self-supervised domain adaptation approach (MSDA) is presented to address the domain shift problem. MSDA employs two novel self-supervised learning strategies to adapt the pose estimator to new domains without requiring annotated pose labels, thereby bridging the performance gap caused by domain shift. Lastly, two methods (OVE6D and GS-Pose) are pro-posed to tackle the challenge of model generalizability. OVE6D is a depth-based approach that utilizes a lightweight Convolutional Neural Network (CNN) model to encode object viewpoints into global feature descriptors for template matching. GS-Pose is an RGB-based method that does not require 3D computer-aided design (CAD) models. GS-Pose integrates 2D object detection, 6D pose initialization, and refinement into a unified framework, streamlining the process of estimating 6D poses of previously unseen objects in practical scenarios.
Firstly, this dissertation introduces a deep neural network model (SC6D) designed to address pose ambiguity. SC6D maps RGB images and object poses into a shared latent space, enabling the implicit handling of ambiguous poses through learned representations. Building on SC6D, a monocular self-supervised domain adaptation approach (MSDA) is presented to address the domain shift problem. MSDA employs two novel self-supervised learning strategies to adapt the pose estimator to new domains without requiring annotated pose labels, thereby bridging the performance gap caused by domain shift. Lastly, two methods (OVE6D and GS-Pose) are pro-posed to tackle the challenge of model generalizability. OVE6D is a depth-based approach that utilizes a lightweight Convolutional Neural Network (CNN) model to encode object viewpoints into global feature descriptors for template matching. GS-Pose is an RGB-based method that does not require 3D computer-aided design (CAD) models. GS-Pose integrates 2D object detection, 6D pose initialization, and refinement into a unified framework, streamlining the process of estimating 6D poses of previously unseen objects in practical scenarios.
Kokoelmat
- Väitöskirjat [5338]
