Workshop Details
TRIPODS Data Science Boot Camp Summer 2021
- Start Date: June 8, 2021
- End Date: June 12, 2021
- Event Start Time: 11:00 AM
- Event End Time: 12:30 PM
- Organizers: Ying Hung | Matthew Stone | Peng Zhang | Hoa Wang
- Location: Online Event
-
Machine Learning Research: Making the most of your data
This boot camp will run June 7-11, 2021 at 11:00 am – 12:30 pm each day.
Machine learning is a powerful tool for making predictions that you can use to make sense of scientific data sets and to develop more flexible and efficient solutions to engineering problems. In many cases, the limiting factor in using machine learning is the difficulty of obtaining data points that you can use to train models. In real applications, data points may take a lot of time to collect (because of the physical or computational processes needed to produce the data), a lot of money to collect (because they involve human effort to collect or annotate), or may simply involve rare events that don’t happen very often even in large data sets. Understanding the problems and research in dealing with limited data is therefore key for using machine learning techniques effectively.
This tutorial will explain some theoretical, computational and practical issues involved in these problems, through the lens of a number of Rutgers faculty who do research in the area. Where possible, sessions will include an interactive component, and will get to experiment with hands-on resources (e.g., python notebooks) that illustrate the problems and solutions discussed. The sessions will be held over zoom.
Registration is free and open but required: https://go.rutgers.edu/2zx0762w
Schedule:Monday, June 7. Matthew Stone (Rutgers, CS)
- Overview. Machine learning and limited data
- Building regression models and the fundamental bias/variance tradeoff.
Tuesday, June 8. Ying Hung (Rutgers, Statistics)
- Experiment design, Part 1.
- Understanding what data you need to improve a model and how to get it.
Wednesday, June 9. Matthew Stone (Rutgers, CS)
- Validating models.
- Visualizing models and understanding the effects of chance in performance.
Thursday, June 10. Peng Zhang (Yale, joins Rutgers CS in fall)
- Experiment design, Part 2.
- Designing data sets to mitigate the effects of chance and improve noisy estimates.
Friday, June 11. Hao Wang (Rutgers, CS)
- Handling imbalanced classes.
- Dealing with rare events with deep learning.
Our goal is that these examples will give you some ideas, programming models and design patterns to support your own data-driven research this summer.
Because everyone is participating remotely, it makes sense to work in the cloud. We’ll be using google colab to explore our data interactively. Things will go smoother on Monday if you take a moment to familiarize yourself with roughly how google colab works: https://colab.research.google.com/notebooks/intro.ipynb
Google colab is nicely integrated with google drive, which is an easy way to share notebooks, data sets and other resources for the class. Materials will be available in advance of each lecture. Here is the link that you can use to obtain materials: https://go.rutgers.edu/97hooha1
When the materials are available, you’ll want to copy this directory into your own google drive, and you’ll then be able to access all the files from within colab.
For more tips on data driven research, check out last summer's boot camp, which focused on visualizations and text data: http://robotics.cs.rutgers.edu/data-inspire/boot-camp/
- Restrictions: Registered Participants
- Audiences: Graduate Students | Undergraduate Students
