Embodied perception & mapping
Multimodal scene understanding for navigation and manipulation: geometry-aware representations for localization, mapping, and decision-making under real sensing constraints.
Robotics · Computer Vision · Machine Learning
Postdoctoral Research Fellow
Australian Institute for Machine Learning, The University of Adelaide
I build perception and reasoning systems that let robots understand, map, and act in the messy geometry of the real world — from structure-aware SLAM to generalizable Gaussian splatting and vision-language reasoning for robot introspection.
I’m a robotics and computer vision researcher working on visual mapping and localization, 3D reconstruction, and multimodal scene understanding for embodied systems. My aim is to give robots representations of the world that are geometrically faithful, semantically meaningful, and robust enough to survive deployment.
I completed my PhD at the University of Adelaide under Prof. Ian Reid at the Australian Centre for Robotic Vision, on real-time, structure-aware, object-centric semantic visual SLAM — integrating geometric estimation with object- and scene-level priors for reliable operation in complex environments. Between my PhD and my current role I spent several years in industry, deploying real-time SLAM, mapping, and situational-awareness systems on field and warehouse robots. That experience still shapes how I work: runtime, sensing limits, and failure modes are first-class concerns, not afterthoughts.
Multimodal scene understanding for navigation and manipulation: geometry-aware representations for localization, mapping, and decision-making under real sensing constraints.
Learning-based reconstruction and neural rendering — Gaussian splatting and NeRF — with an emphasis on geometric consistency and generalization beyond the training distribution.
Using VLMs over structured visual evidence for robot introspection: detecting and localizing failures in long-horizon tasks, producing grounded explanations, and proposing corrections.
Selected work, newest first. A complete list is available on Google Scholar.
Injects surface-normal and depth priors into pose-free generalizable Gaussian splatting built on geometric foundation models, improving reconstruction geometry in textureless and complex indoor scenes.
A training-free front-end that compresses long robot-execution videos into compact, interpretable evidence, substantially improving VLM failure detection, identification, and localization.
An RGB-only topometric navigation pipeline combining global topological planning with local metric trajectory control for zero-shot, long-horizon navigation without 3D maps.
Uses sensor pose as a supervisory signal to align camera and LiDAR bird’s-eye-view embeddings, outperforming fully-supervised BEV map segmentation with a fraction of the annotated data.
A segment-level topological representation of the environment that supports zero-shot navigation driven by open-vocabulary language queries.
Integrates deep-learned object detections and CNN shape priors as quadrics into a monocular SLAM factor graph, refining object geometry while retaining real-time performance.
Represents generic objects as quadrics and scene structure as planes in a unified monocular SLAM map, with shape and plane–plane constraints tying semantics to geometry.
2022 – present
Australian Institute for Machine Learning (AIML), The University of Adelaide
2020 – 2022
2019 – 2020
KikTech Robotics — startup environment
2016 – 2019
Australian Centre for Robotic Vision, The University of Adelaide
Ten years of university teaching across programming, algorithms, and AI — from first-year foundations to postgraduate coursework and research supervision.
2024 – 2026
Kaplan Business School
Owned curriculum, assessment, and delivery for postgraduate IT subjects; handled grade moderation and academic liaison.
2022
The University of Adelaide
Coordinated Algorithm & Data Structure Analysis (Semester 2, 2022), delivering lectures and managing the tutoring team.
2017 – 2026
The University of Adelaide College & Kaplan International College
Developed and delivered modules across computing and electronics; recognised in 2018 for teaching excellence on the basis of SELT results.
M. T. Aamir · Master’s project, 2024
Integrated large language models with 2D LiDAR, YOLO-based detection, and NAV2 for spatial reasoning in dynamic indoor navigation.
X. Hu · Master’s project, 2024
Built a gesture recognition pipeline using deep learning for real-time 3D motion tracking and feedback across varied environments.
Long before SLAM, robot soccer taught me how real-time autonomy actually behaves. In 2008 I co-founded the AI group of the Omid Robotics Team, competing in the Small Size League of RoboCup — a hybrid centralised/distributed multi-agent problem where perception, planning, and control all run inside a few milliseconds.
Semantics and Objects in SLAM
Computer vision & machine learning Reviewer for CVPR, NeurIPS, ICCV, ECCV, and BMVC since 2023 — Outstanding Reviewer, ECCV 2024.
Robotics Reviewer for ICRA, IROS, IEEE Transactions on Robotics (T-RO), and IEEE Robotics and Automation Letters (RA-L) since 2019.