Computer vision · Multimodal learning

Simon Jenni

I develop vision–language systems for creative products—from search and retrieval to content authenticity and intelligent editing.

R / 01

Vision–language models

Multimodal representations for visual search, retrieval, aesthetics, recommendation, and spatial understanding.

Embeddings · VLMs · Retrieval

R / 02

Visual fingerprinting

Robust image and video matching for provenance, deduplication, asset management, and moderation at scale.

Matching · Provenance · Trust

R / 03

Self-supervised learning

Learning transferable representations from the natural structure of images, video, and audio.

Video · Audio · Contrastive learning

2026
2025

Improving Large Vision and Language Models by Learning from a Panel of Peers

J. Hernandez, J. Shi, S. Jenni, V. Ordonez, K. Kafle

ICCV

MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling

S. Khosla, A. Tiwari, K. Kafle, S. Jenni, H. Zhao, J. Collomosse, J. Shi

ACL
2024

Building Vision-Language Models on Solid Foundations with Masked Distillation

S. Sameni, K. Kafle, H. Tan, S. Jenni

CVPR

Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets

I. Dave, F. C. Heilbron, M. Shah, S. Jenni

ECCVOral

FINEMATCH: Aspect-Based Fine-Grained Image and Text Mismatch Detection

H. Hua, J. Shi, K. Kafle, S. Jenni, D. Zhang, J. Collomosse, S. Cohen, J. Luo

ECCV

Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models

G. Kwon, S. Jenni, D. Li, J. Lee, J. Ye, F. Caba Heilbron

CVPR
Earlier

Audio-Visual Contrastive Learning with Temporal Self-Supervision

S. Jenni, A. Black, J. Collomosse

AAAI ’23

Time-Equivariant Contrastive Video Representation Learning

S. Jenni, H. Jin

ICCV ’21Oral

Video Representation Learning by Recognizing Temporal Transformations

S. Jenni, G. Meishvili, P. Favaro

ECCV ’20

Self-Supervised Feature Learning by Learning to Spot Artifacts

S. Jenni, P. Favaro

CVPR ’18Spotlight

Recent

Jun 2026Gen2Balance accepted to ECCV 2026.
Jun 2026RetouchIQ presented at CVPR 2026.
Apr 2026Seeing Through Words presented at ICLR 2026.
2025Five papers across CVPR, ICCV, NeurIPS, and ACL.
Jan 2025Promoted to Senior Research Scientist at Adobe Research.
Oct 2024Video fingerprinting featured in Project KnowHow at Adobe MAX.

Recognition

2022Faculty Prize for best PhD dissertation, University of Bern.
2021Best Paper Award, CVMP.
2019–22Outstanding Reviewer, CVPR and ECCV.
2018Best Poster Award, PRAIRIE/MIAI AI Summer School.
2017Best Master Thesis, Joint Alumni Association in Computer Science.