• Start Date: March 5, 2025
  • Event Start Time: 11:00 AM
  • Event End Time: 12:00 PM
  • Seminar Series: Theoretical Computer Science Seminar
  • Presenter(s): Vincent Cohen-Addad - Google
  • Event Location: Conference Room 301 | Rutgers University | CoRE Building | 96 Frelinghuysen Road
  • Event Additional Info: <p>See:&nbsp;<a href="https://theory.cs.rutgers.edu/theory_seminar">https://theory.cs.rutgers.edu/theory_seminar</a></p>
  • Presentation Type: Stand Alone Presentation
  • Abstract:

    The scale of modern machine learning models and data has made data selection a central problem. In this talk, we focus on the problem of finding the best representative subset of a dataset to train a machine learning model. We provide a new data selection approach based on ð‘˜-means clustering and sensitivity sampling.

    Assuming embedding representation of the data and that the model loss is Hölder continuous with respect to these embeddings, we prove that our new approach allows to select a set of \`\`typical'' 1/𝜖2 elements whose average loss corresponds to the average loss of the whole dataset, up to a multiplicative (1±ðœ–) factor and an additive ðœ–𝜆Φ𝑘, where Φ𝑘 represents the ð‘˜-means cost for the input data and ðœ† is the Hölder constant. We furthermore demonstrate the performance and scalability of our approach on fine-tuning foundation models and show that it outperforms state-of-the-art methods.

    We also show that our sampling strategy can be used to define new sampling scores for regression, leading to a new active learning strategy that is comparatively simpler and faster than previous ones like leverage score.

    Based on several papers that appeared at FOCS'24 and ICML'24, joint work with Kyriakos Axiotis, Nikhil Bansal, Monika Henzinger, Sammy Jerome, Vahab Mirrokni, Milind Prabhu, David Saulpic, Chris Schwiegelshohn, David Woodruff, Michael Wunder