Text-and-audio methods

Date:

Tuesday, 4 February, 2025 - 13:00 to 14:00

Speaker:

Cătălina Cangea, ex-Google DeepMind

Venue:

Lecture Theatre 2, Computer Laboratory, William Gates Building

This talk supports the R255 Advanced Topics in Machine Learning module on Multimodal Learning and provides a bird’s eye view of the rapidly evolving text-audio landscape, with a focus on music as a primary example of audio data. I will first present types of tasks that exist in this space, then discuss data curation challenges and follow with an overview of some existing retrieval and generation methods, including a quick primer on diffusion models. Finally, I will describe current evaluation metrics and their limitations.

Seminar series:

Artificial Intelligence Research Group Talks

View on talks.cam

Calendar

Upcoming seminars

About the department

Social media

Study at Cambridge

About the University

Research at Cambridge