Incoming MSc Student · University of Bonn

Gagneet Singh

AI Researcher & Incoming MSc Student at the University of Bonn

I am an AI researcher working at the intersection of computer vision, medical image analysis, and multimodal learning. My research focuses on developing deep learning systems that can understand complex visual and linguistic information, with a particular interest in vision-language models, representation learning, and the interpretability of large multimodal models. I will be joining the University of Bonn as a Master’s student in Winter Semester 2026, where I aim to deepen my research in these areas and contribute to high-impact work in artificial intelligence.

Computer VisionMedical AIMultimodal Learning
Research Interests

Areas of Focus

My research spans several interconnected areas of artificial intelligence, from visual understanding to multimodal reasoning.

Computer Vision

Developing deep learning systems that can understand, analyze, and interpret visual information across diverse domains.

Medical Image Analysis

Applying AI to medical imaging problems including segmentation, classification, diagnosis, and clinical decision support.

Vision-Language Models

Investigating multimodal models that bridge visual and linguistic understanding for richer AI capabilities.

Multimodal AI

Building systems capable of learning from multiple modalities such as images, text, and structured data.

Large Language Models

Exploring language models, instruction tuning, retrieval, and the foundations of language understanding.

Machine Learning

Developing robust and generalizable machine learning systems with strong theoretical grounding.

Research Overview

A Unified Vision for Intelligent Systems

My research focuses on developing intelligent learning systems capable of understanding complex visual and multimodal information. I am particularly interested in applying deep learning and representation learning to challenging problems in computer vision and medical image analysis, while also exploring the capabilities, limitations, and interpretability of large multimodal models.

Vision-Language ModelsMedical Image AnalysisRepresentation LearningLarge Language Models
Learn more about my research

Research Profile

Current Focus
Computer Vision
Medical AI
Multimodal Learning
Upcoming
MSc at the University of Bonn
Research Interests
Vision-Language Models
Medical Image Analysis
Representation Learning
Large Language Models
Featured Research

Selected Projects

A selection of ongoing and completed research projects spanning medical imaging, vision-language models, and multimodal AI.

Medical Image Analysis

Automated Cervical Vertebral Maturation Staging

Completed

Automated assessment of cervical vertebral maturation stages using deep learning and ordinal learning approaches for clinical decision support in orthodontics.

Panjab University / PGIMER

20242025

Vision-Language Models

Vision-Language Model Robustness and Interpretability

In Progress

Research investigating the failure modes, robustness, and internal mechanisms of vision-language models, including noise robustness and mechanistic interpretability.

MBZUAI

2025

Multimodal AI

Medical Imaging and Radiology Report Generation

In Progress

Multimodal AI systems leveraging vision-language models for medical imaging and automated radiology report generation.

MBZUAI

2025

Publications

Selected Publications

Recent and forthcoming publications in computer vision, medical image analysis, and multimodal learning.

AcceptedConferenceFeatured

Automated Cervical Vertebral Maturation Staging from a Continuous Perspective

Gagneet Singh, Author Two, Author Three
IEEE ICCCNT 2025·2025

An automated approach for assessing cervical vertebral maturation stages using deep learning and ordinal learning, enabling continuous-stage prediction from lateral cephalometric radiographs.

Medical Image AnalysisDeep LearningOrdinal Regression
Oral PresentationWorkshopFeatured

On the Robustness and Interpretability of Vision-Language Models

Gagneet Singh, Author Four, Author Five
NeurIPS 2025 Workshop·2025

An investigation into the failure modes, robustness, and internal mechanisms of vision-language models, with a focus on noise robustness and mechanistic interpretability.

Vision-Language ModelsInterpretabilityRobustness