Prof.
Stefan Roth, PhD
Technische Universität Darmstadt
Visual Inference
Hochschulstraße 10
64289 Darmstadt
Short info
My research focuses on machine learning approaches to understanding and analyzing digital images and videos. My lab and I develop new deep learning models and methods for scene understanding, motion estimation, image editing and synthesis, video analysis, image restoration, and more. We aim to make such approaches robust to a broad range of real-world conditions. To that end, we are incorporating inductive biases through combining deep learning with classical models, endowing deep networks with explicit representations of uncertainty, or adapting pre-trained models to changing test-time circumstances. We also aim to reduce the dependency on labeled data by developing semi-supervised and self-supervised learning pipelines.
Open Science
Boosting Unsupervised Semantic Segmentation with Principal Mask Proposals.
arXiv preprint2404.16818.
Beyond Accuracy: What Matters in Designing Well-Behaved Models?.
arXiv preprint arXiv: 2503.17110
Activation Subspaces for Out-of-Distribution Detection.
arXiv preprint arXiv: 2508.21695.
Articles
Semantic Self-adaptation: Enhancing Generalization with a Single Sample. Transactions on Machine Learning Research (TMLR).
Boosting Omnidirectional Stereo Matching with a Pre-trained Depth Foundation Model.
Proc of the IEEE/RSJ International Conference on Intelligent Robots and Systems.
Boosting Omnidirectional Stereo Matching with a Pre-trained Depth Foundation Model.
IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS).
Motion-Refined DINOSAUR for Unsupervised Multi-Object Discovery.
IEEE International Conference on Computer Vision (ICCV) Workshops.
Scene-Centric Unsupervised Panoptic Segmentation.
Proceedings of the Computer Vision and Pattern Recognition Conference, 24485-24495.
Disentangling Polysemantic Channels in Convolutional Neural Networks.
Proceedings of the Computer Vision and Pattern Recognition Conference, 4799-4803.
Fast axiomatic attribution for neural networks.
Advances in Neural Information Processing Systems, 34(2), 19513-19524.
Funnybirds: A synthetic vision dataset for a part-based analysis of explainable ai methods.
Proceedings of the IEEE/CVF International Conference on Computer Vision, 3981-3991.
Benchmarking the attribution quality of vision models.
Advances in Neural Information Processing Systems, 37, 97928-97947.
Content-adaptive downsampling in convolutional neural networks.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4544-4553.
Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion.
Proc. of the Twentieth IEEE International Conference on Computer Vision.
Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion.
IEEE International Conference on Computer Vision (ICCV)