• Start Date: June 19, 2008
  • Event Start Time: 12:00 PM
  • Event End Time: 1:00 PM
  • Organizers: Christine Agnese
  • Seminar Series: REU Seminar
  • Presenter(s): Rebecka Jornsten - Rutgers University
  • Event Location: DIMACS Seminar room
  • Abstract: Clustering is an exploratory technique used to discover structure in high-dimensional data. With the data deluge in the biological sciences (so-called high-throughput biology), clustering has become an ever-increasingly popular approach to summarize data. Clustering can be used to detect problems in the data: do experimental batches group together? do experimental data collected on certain days group together? Such clustering results indicate that the experimental design is flawed and systematic biases are present. clustering can also be used to discover biologically relevant groups in the data: do experimental data collected from a subset of patients group together? If so, perhaps these patients share some traits that can provide insight into a disease state or suggest a new path to treatment. In this talk, I will review some statistical methods and concepts for clustering and illustrate the results on a particular type of high-throughput biological data: microarray data.