Accurate 6D object pose estimation is vital for robotics, AR, and autonomous systems but remains difficult in open-set scenarios with unseen objects. This thesis introduces three contributions: MatchU, a multimodal RGB-D descriptor fusion framework; TTAPose, a self-supervised test-time adaptation method; and RayPose, a diffusion-based RGB-only approach. Together they advance robust, generalizable pose estimation for real-world deployment.
Übersetzte Kurzfassung:
Die präzise 6D-Objekt-Pose-Schätzung ist entscheidend für Robotik, AR und autonome Systeme, bleibt jedoch in offenen Szenarien mit unbekannten Objekten herausfordernd. Diese Arbeit stellt drei Beiträge vor: MatchU, ein multimodales RGB-D-Framework; TTAPose, eine selbstüberwachte Testzeit-Adaptation; und RayPose, ein diffusionsbasiertes RGB-Only-Verfahren. Zusammen fördern sie eine robuste und generalisierbare Pose-Schätzung für reale Anwendungen.