Vision–language models
Multimodal representations for visual search, retrieval, aesthetics, recommendation, and spatial understanding.
Embeddings · VLMs · Retrieval
Computer vision · Multimodal learning
I develop vision–language systems for creative products—from search and retrieval to content authenticity and intelligent editing.
Learning representations that connect pixels, language, and creative intent.
Multimodal representations for visual search, retrieval, aesthetics, recommendation, and spatial understanding.
Embeddings · VLMs · Retrieval
Robust image and video matching for provenance, deduplication, asset management, and moderation at scale.
Matching · Provenance · Trust
Learning transferable representations from the natural structure of images, video, and audio.
Video · Audio · Contrastive learning
Research translated into tools used by creative professionals.
Recent and selected publications. Titles link to the paper or official record.