Dishani Lahiri

I am a Machine Learning Engineer 2 at Snap Inc. in New York City, working on post-training efficient generative models for image and video.

I completed my Master of Science in Computer Vision (MSCV) in the Robotics Institute at Carnegie Mellon University in Dec 2023, where I was a Research Assistant at KLab, advised by Prof. Kris Kitani, working on visual-inertial state estimation for Meta Aria glasses.

During my summer internship at Slingshot AI, I worked on fine-tuning personalized text-to-image diffusion models and LLaMA2-7B for personalized text style transfer.

Before that, I worked at Samsung R&D Institute, Bangalore, where I helped develop and deploy AI Night mode and the Expert RAW app for Samsung's flagship devices.

I completed my undergraduate studies in ECE at DTU in 2019, advised by Prof. S. Indu and Prof. D.K. Vishwakarma.

Email  /  Google Scholar  /  LinkedIn  /  Github

profile photo
sym

Snap Inc.
ML Engineer 2
Feb. 24 - Present

sym

CMU
MS in Computer Vision
Aug. 22 - Dec. 23

sym

KLab@CMU
Research Assistant
Jan. 22 - Dec. 23

sym

Slingshot AI
ML Research Intern
May 23 - Aug. 23

sym

Samsung R&D Institute
Senior CV Engineer
June 19 - July 22
Software Engineer Intern
May 18 - July 18

sym

DTU, Delhi
B.Tech. ECE
Aug. 15 - May 19

Projects & Publications

I'm interested in computer vision, natural language processing, and machine learning, especially in building personalized multi-modal solutions for edge devices.

S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
CVPR, 2026
paper | poster

A streaming sandwich diffusion transformer for efficient, high-fidelity video generation on mobile devices. S2DiT uses a mixture of LinConv Hybrid Attention and Stride Self-Attention, a sandwich architecture found via budget-aware dynamic programming, and a 2-in-1 distillation framework from large teacher models, achieving quality on par with state-of-the-art server video models while streaming at over 10 FPS on an iPhone.

SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices
arXiv preprint, 2026
paper

A diffusion transformer for efficient, high-fidelity image generation on edge devices, combining an adaptive global-local sparse attention mechanism, an elastic training framework for dynamic adjustment across hardware targets, and a distillation pipeline enabling high-fidelity, low-latency (4-step) generation suitable for real-time on-device use.

SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
CVPR, 2025 (Highlight)
paper | project page

An efficient text-to-image diffusion model for mobile devices, combining architectural design choices, cross-architecture knowledge distillation from larger models, and adversarial few-step guidance. SnapGen generates 1024x1024 images on a mobile device in about 1.4s with only 379M parameters, outperforming significantly larger models on standard benchmarks.

S2RF: Semantically Stylized Radiance Fields
Dishani Lahiri*, Neeraj Panse*, Moneish Kumar*
ICCV, 2023 Workshop on AI for 3D Content Creation
paper | code | webpage

We present our method for transferring style from any arbitrary image(s) to object(s) within a 3D scene. Our primary objective is to offer more control in 3D scene stylization, facilitating the creation of customizable and stylized scene images from arbitrary viewpoints. To achieve this, we propose a novel approach that incorporates nearest neighborhood-based loss, allowing for flexible 3D scene reconstruction while effectively capturing intricate style details and ensuring multi-view consistency.

Abnormal human action recognition using average energy images
Dishani Lahiri*, Chhavi Dhiman, Dinesh Kumar Vishwakarma
IEEE, 2017 Conference on Information and Communication Technology (CICT)
paper

We propose a solution to detect abnormal human actions in the image using Histogram of Oriented Gradients (HoG) as the feature descriptor, Principal Component Analysis (PCA) as the dimensionality-reduction technique, and Support Vector Machine as the ML tool for classification. We also release a dataset for abnormal human activities of fainting, headache, and chest pain.

Teaching Experience
  • Advanced Computer Vision, CMU (TA) | Instructor: Prof. David Held | Fall 2023
    This is a new PhD-level course wherein I am involved in preparing and improving the assignments, maintaining the course website, holding Office Hours, and helping students with the theory and code of concepts covered throughout the course.
  • Machine Learning, CMU (TA) | Instructor: Prof. Matt Gormley | Spring 2023
    Preparing and suggesting exam and assignment problems, and material in order to make the course more effective. Holding recitations and office hours for students.
Awards and Recognition
  • Winner (most creative use of Github), HackCMU : Awarded for our project, How Do I Look?, using image-to-text and Large Language Models to generate suggestions for attires based on the event
  • Samsung Excellence Award (earlier Samsung Citizen Award), Advanced Development Category : Company-wide Award to recognize major contributions towards the R&D in Night Mode for S21 Flagship series
  • Standout Performer in Advanced R&D Work : Succeeded in being 1 out of 100 people in Camera Systems Group to receive this award for constant exceptional efforts towards research and implementation
  • Samsung Citizen Award, Group Excellence Category : Company-wide Group award to recognize major contributions towards the development of camera usecases in A71-5G device, the first device with SM7250 chipset
  • Standout Performer in Advanced R&D Work : Succeeded in being 1 out of 100 people in Camera Systems Group to receive this award for constant exceptional efforts towards research and implementation
  • 1H-2020 Project Incentives : Succeeded in being 1 in 2 out of 100 people in Camera Systems Group to receive the incentive in lieu of exceptional performance in critical projects
  • Appreciation letter from HRD Ministry of India : For being in top 0.1 percentile scorers in 12th class CBSE examination. HRD Ministry is the Government of India Body formulates the National Policy of Education