Portrait of Babak Ehteshami Bejnordi

Babak Ehteshami Bejnordi

Principal Research Scientist at NVIDIA in Zurich, working on the Nemotron team and developing efficient architectures for large language models.

I am interested in building capable, efficient AI models and advancing the field through open and reproducible research.

Previously, I was a Senior Staff Research Scientist and Manager at Qualcomm AI Research in Amsterdam, where I led a team focused on efficient LLM architectures. My work covered mixture-of-experts models, efficient and latent reasoning, on-device LLM deployment, computer vision, multi-task learning, and continual learning. I also organized the Qualcomm Innovation Fellowship Program in Europe from 2019 to 2023.

I obtained my PhD at the Diagnostic Image Analysis Group, Radboud University, where I developed machine-learning algorithms for breast cancer diagnostics and organized the CAMELYON16 challenge.

From June to November 2016, I was a visiting researcher at Harvard University, studying tumor-associated stroma as a prognostic biomarker in breast cancer with collaborators from Harvard, NIH, and Mayo Clinic.

NVIDIA Β· Zurich, Switzerland

Research updates

Latest research

View all research
Visualization of Dirichlet-Prior Shaping for expert specialization
ICML 2026

Dirichlet-Prior Shaping

Guiding expert specialization in upcycled mixture-of-experts.

Read paper
Efficient reasoning on edge devices project overview
Technical report 2026

Reasoning on the Edge

Reasoning in small LLMs with LoRA adapters, supervised fine-tuning, and reinforcement learning.

Read paper
KaVa compressed KV-cache distillation architecture
ICLR 2026

Latent Reasoning

Distilling knowledge from a compressed teacher KV-cache into a latent-reasoning student.

Read paper
Cache-MoE expert caching architecture for mobile inference
TMLR 2025

Cache-MoE

Efficient mixture-of-experts inference on mobile devices with limited DRAM.

Read paper
READ-ME router-decoupled mixture-of-experts architecture
NeurIPS 2024

Refactor LLM into MoE

Refactorizing LLMs as router-decoupled mixture-of-experts with system co-design.

Read paper
LLM-to-SLM fast autoregressive decoding method
ICML workshop 2024

LLM-to-SLM

Combining large and small language models for fast autoregressive decoding.

Read paper
InterroGate representation sharing and specialization method
BMVC 2024

InterroGate for MTL

Learning to share, specialize, and prune representations for multi-task learning.

Read paper
Scalarization method for multi-task and multi-domain learning
NeurIPS 2023

Scalarization for MTL

Scalarization for multi-task and multi-domain learning at scale.

Read paper