Workshop Details
IBM/DIMACS/DATA-INSPIRE Workshop on Bridging Game Theory and Machine Learning for Multi-party Decision Making
- Start Date: October 28, 2022
- End Date: October 29, 2022
- Event Start Time: 8:30 AM
- Event End Time: 3:30 PM
- Organizers: Segev Wasserkrug | David Pennock | Tamra Carpenter
- Location: Rutgers University Inn and Conference Center | Rutgers University | 178 Ryders Lane
-
Many real-world decision-making situations involve the decisions of multiple parties. Typically, in such situations each party wants to optimize its own objectives, even knowing that their actions affect the objectives of others and vice versa. Game theory is the mathematical science intended to model such situations. Yet many real decision-making scenarios are far too complex for traditional game theory to provide prescriptive recommendations; for example:
- Business networks: settings where multiple, possibly independent, enterprises cooperate in order to obtain more value as a network than they can as individual firms. An example of such a network is a retail supply chain, consisting of supermarket chains, distributors, farmers, etc. Such settings result in an interesting combination of cooperation and competition: firms cooperate to increase the network value, while simultaneously competing on the division of this value.
- Cloud and hybrid cloud computer systems: Analyzing games in which the players include, but are not limited to: cloud providers, enterprise infrastructure teams, companies using the IT infrastructure to provide services to their customers, and the end customers themselves. Problems in this space include contracting, pricing decisions, infrastructure usage decisions, and workload placement.
- Cybersecurity: optimizing decisions for improving cybersecurity by analyzing such settings as adversarial games between defenders and attackers.
Traditional game-theoretic assumptions, including unbounded rationality, unlimited computation, common knowledge, and common priors, often don’t apply in these settings. Algorithmic game theory relaxes some of these assumptions, examining the behavior of parties and mechanisms with bounded computational resources, while simultaneously trying to provide clear and simple mechanisms for the participants to follow. Still, multiple solution concepts with multiple equilibria dilute the predictions that (algorithmic) game theory can produce, and there are many additional real-world scenarios in which computationally efficient and clear mechanisms have yet to be created. Finally, multiagent reinforcement learning (MARL) and other machine learning techniques offer a more prescriptive approach to teach agents how to behave in multi-party settings. These techniques sometimes work quite well in practice, and in vastly more complex settings than traditional game theory can handle, yet less is known about their theoretical convergence properties and performance guarantees.
It is therefore the goal of this workshop to bring together researchers from both industry and academia in the domains of game theory, algorithmic game theory, multi agent reinforcement learning, and learning in game theory to understand and study the problems in which multi-party decisions are required, create joint awareness of the current state-of-the-art in both industry and academia, including relevant software tools and platforms, and seed collaborations both in the integration and scientific advancement of these techniques, as well as their application to real world use cases such as the ones described above.
Topics of interest include, but are not limited to:- Real-world business scenarios in which such techniques are required, as well as the theoretical and practical gaps inhibiting their application.
- Solution concepts relevant to such real-world use cases and their efficient computation.
- Building upon commonly used knowledge frameworks such as common knowledge priors and knowledge hierarchies to define knowledge frameworks appropriate for real-world decision-making settings.
- Algorithms enabling agents to learn how to best behave in such settings using algorithms such as MARL and regret minimization, based on sound scientific principles.
- Research on how game theory, algorithmic game theory and MARL algorithms can be successfully applied to scenarios such as cloud computing, supply chain networks and security.
We are grateful to IBM Research for its generous support of this event.
-
Workshop Additional Information
Parking Information
Visitors may park in Lots 74A, 76, and 82. They must use the link below to register for their event. Until this process is completed their vehicles are not registered and they may receive a citation. Special event parking and special event permits are only for visitors to the University, which does not include free metered parking. Faculty, Staff, and Students must park only in lots they are authorized to park in and should not register using this link.
Directions to parking lots can be found at http://maps.rutgers.edu
-
Thursday, October 27, 2022
Workshop Talks
8:30 AM – 9:00 AMWelcome from Organizers
8:50 AM – 10:00 AMKeynote: The State of Representing and Solving Games
Tuomas Sandholm - Carnegie Mellon University
9:00 AM – 10:30 AMBreakfast is available beginning at 8:30
Speaker Bio: Tuomas Sandholm is Angel Jordan University Professor of Computer Science at Carnegie Mellon University and a serial entrepreneur. His research focuses on the convergence of artificial intelligence, economics, and operations research. He is Co-Director of CMU AI. He is the Founder and Director of the Electronic Marketplaces Laboratory. He has published over 500 peer-reviewed papers, holds 25 US patents, and his h-index is 91. In addition to his main appointment in the Computer Science Department, he holds appointments in the Machine Learning Department, Ph.D. Program in Algorithms, Combinatorics, and Optimization (ACO), and CMU/UPitt Joint Ph.D. Program in Computational Biology.
He has built optimization-powered electronic marketplaces since 1989, and has fielded several of his systems. In parallel with his academic career, he was Founder, Chairman, first CEO, and CTO/Chief Scientist of CombineNet, Inc. from 1997 until its acquisition in 2010. During this period the company commercialized over 800 of the world's largest-scale generalized combinatorial multi-attribute auctions, with over $60 billion in total spend and over $6 billion in generated savings.
Since 2010, his algorithms have been running the national kidney exchange for the United Network for Organ Sharing, where they autonomously make the kidney exchange transplant plan for 80% of U.S. transplant centers together each week. He also co-invented never-ending altruist-donor-initiated chains and his algorithms created the first such chain. Such chains have become the main modality of kidney exchange worldwide and have led to around 10,000 life-saving transplants. He invented liver lobe and multi-organ exchanges, and the first liver-kidney swap took place in 2019.
Sandholm has developed the leading algorithms for several general classes of game with his students. The team that he leads is the multi-time world champion in computer heads-up no-limit Texas hold’em, which is the main benchmark and decades-open challenge problem for testing application-independent algorithms for solving imperfect-information games. Their AI Libratus became the first and only AI to beat top humans at that game. Then their AI Pluribus became the first and only AI to beat top humans at the multi-player game. That is the first superhuman milestone in any game beyond two-player zero-sum games. He is Founder and CEO of Strategic Machine, Inc., which provides solutions for strategic reasoning under imperfect information in a broad set of applications ranging from poker to other recreational games to business strategy, negotiation, strategic pricing, finance, cybersecurity, physical security, auctions, political campaigns, and medical treatment planning. He is also Founder and CEO of Strategy Robot, Inc., which focuses on defense, intelligence, and other government applications. Strategy Robot has built software products based on game-theoretic reasoning technology for portfolio planning, base defense, missile defense, and course-of-action generation.
He is Founder and CEO of Optimized Markets, Inc., which is bringing a new optimization-powered paradigm to advertising campaign sales, pricing, and scheduling - in TV (linear and nonlinear), Internet display, streaming (video and audio), mobile, game, and cross-media advertising.
He served as the redesign consultant of Baidu’s sponsored search auctions and display advertising markets; within two years Baidu’s market cap increased 5x to $50 billion due to doubled monetization per user. He has served as consultant, advisor, or board member for Yahoo!, Google, Chicago Board Options Exchange, swap.com, Granata Decision Systems (now part of Google), Rare Crowds (now part of Media Math), and others.
He holds a Ph.D. and M.S. in computer science and a Dipl. Eng. (M.S. with B.S. included) with distinction in Industrial Engineering and Management Science. Among his many honors are the Minsky Medal, IJCAI McCarthy Award, AAAI/IAAI Engelmore Award, IJCAI Computers and Thought Award, inaugural ACM Autonomous Agents Research Award, CMU’s Allen Newell Award for Research Excellence, Sloan Fellowship, NSF Career Award, Carnegie Science Center Award for Excellence, and Edelman Laureateship. He is Fellow of the ACM, AAAI, INFORMS, and AAAS. He holds an honorary doctorate from the University of Zurich.
10:00 AM – 10:40 AMCustomer-Centric Science: The Sponsored Products Example
Muthu Muthukrishnan - Amazon
Science is inspired to study natural, chemical, economic and other systems and has broad set of tools. In this talk, we consider the Science that arises from trying to understand he needs of online customers. We will use the Sponsored Products example which is an advertising solution for customers' needs. We will provide an overview of the problems and tools.
Speaker Bio: Muthu Muthukrishnan is Vice President of Sponsored Products for Amazon Advertising, a self-service advertising solution that drives product discovery and sales on Amazon.com. Prior to joining Amazon in 2018, Muthukrishnan held a variety of research positions, including those at Bell Laboratories, Microsoft, and Google, as well as Professor of Computer Science at Rutgers University. Muthu's scientific interests are in Algorithms, as well as Auctions and Game Theory. Muthu is a Fellow of the ACM and winner of the Imre Simon Test of Time Award for work on count-min sketch.
10:40 AM – 11:00 AMBreak (20 minutes)
11:00 AM – 11:40 AMMulti-agent Learning and Equilibrium (remote presentation)
Bernhard von Stengel - London School of Economics
This project proposes a framework, under development, of using machine learning to study larger and more realistic game-theoretic models. In a dynamic pricing game, firms compete by repeatedly setting a price for a product, where higher prices lead to higher short-term but lower long-term profits. The classic subgame-perfect equilibrium is very competitive, whereas strategic experiments show a tendency of agents to collude. We want to employ small neural nets to learn pricing strategies, with the learning environment given by a mixed equilibrium of existing strategies. Successfully learned strategies are added to the pool and a new equilibrium is computed. This is akin to double-oracle learning in zero-sum games, except that the game is not zero-sum and the resulting equilibrium depends on the learning history. The approach is modular rather than a large simulation, which should allow a better study of the relevant features of the underlying game, its learning mechanism, and the employed equilibrium concept.
Speaker Bio: Bernhard von Stengel, educated in Germany and the US, is Professor of Mathematics at the London School of Economics. He is a mathematical game theorist and an expert on computational and geometric methods for solving games. He chaired the 2016 World Congress of the Game Theory Society, is co-editor of the International Journal of Game Theory, and was Area Editor for Game Theory for Mathematics of Operations Research. His recent textbook "Game Theory Basics" teaches standard and non-standard topics that every game theorist should know.
11:40 AM – 12:20 PMMechanism Learning for Trading Networks
Takayuki Osogami - IBM Research
We study the problem of designing mechanisms for trading networks that satisfy four desired properties: dominant-strategy incentive compatibility, efficiency, weak budget balance (WBB), and individual rationality (IR). Although there exist mechanisms that simultaneously satisfy these properties ex post for combinatorial auctions, we prove the impossibility that such mechanisms do not exist for a broad class of trading networks. We thus propose approaches for computing and learning the mechanisms that satisfy the four properties, in a Bayesian setting, where WBB and IR, respectively, are relaxed to ex ante and interim. For computational and sample efficiency, we introduce several techniques, including game theoretical analysis to reduce the input feature space. We empirically demonstrate that the proposed approaches successfully find the mechanisms with the four properties for those trading networks where the impossibility holds ex post.
Speaker Bio: Takayuki Osogami is a senior technical staff member and the manager of the mathematical sciences group at IBM Research - Tokyo. He is currently leading global research projects on reinforcement learning, industrial applications of reinforcement learning, automation of decision optimization with reinforcement learning, multi-agent reinforcement learning, and integration of learning and game theory. He was a group leader of a governmental project supported by Core Research for Evolutionary Science and Technology, Japan Science and Technology Agency during 2013-2019, where his group developed and applied theory of sequential decision making, human behavior modeling, and neuromorphic computing. He received his Ph.D. in Computer Science from Carnegie Mellon University in August 2005, and a B.Eng. degree in Electronic Engineering from the University of Tokyo in 1998. He was selected as one of the Researchers with Nice Step 2020 by the National Institute of Science and Technology Policy (NISTEP) of the Ministry of Education, Culture, Sports, Science and Technology.
12:20 PM – 1:20 PMLunch (1 hour)
1:20 PM – 2:20 PMKeynote: No-Regret Learning in Extensive-Form Games
Amy Greenwald - Brown University
The convergence of $Phi$-regret-minimization algorithms in self-play to $Phi$-equilibria is well understood in normal-form games (NFGs), where $Phi$ is the set of deviation strategies. This talk investigates the analogous relationship in extensive-form games (EFGs). While the primary choices for $Phi$ in NFGs are internal and external regret, the space of possible deviations in EFGs is much richer. We restrict attention to a class of deviations known as behavioral deviations, inspired by von Stengel and Forges' deviation player, which they introduced when defining extensive-form correlated equilibria (EFCE). We then propose extensive-form regret minimization (EFR), a regret-minimizing learning algorithm whose complexity scales with the complexity of $Phi$, and which converges in self-play to EFCE when $Phi$ is the set of behavioral deviations. Von Stengel and Forges, Zinkevich et al., and Celli et al. all weaken the deviation player in various ways, and then derive corresponding efficient equilibrium-finding algorithms. These weakenings (and others) can be seamlessly encoded into EFR at runtime, by simply defining an appropriate $Phi$. The result is a class of efficient $Phi$-equilibrium finding algorithms for EFGs.
Speaker Bio: Amy Greenwald is Professor of Computer Science at Brown University in Providence, Rhode Island. Greenwald was also a visiting researcher at the Artificial Intelligence Research Center at the Japanese National Institute of Advanced Industrial Science and Technology in Tokyo in 2018-19; a visiting researcher in the Algorithmic Economics Lab at Microsoft Research in New York City in 2015; and a visiting professor at the Erasmus Research Institute of Management in Rotterdam in 2011. She was named a Fulbright Scholar in 2011 (though she declined the award); she was awarded a Sloan Fellowship in 2006; she was nominated for the 2002 Presidential Early Career Award for Scientists and Engineers; and she was named one of the Computing Research Association's Digital Government Fellows in 2001. Before joining Brown University, she worked for a short time as a post-doc at IBM's T.J. Watson Research Center, where her "Shopbots and Pricebots" paper was named Best Paper at IBM Research in 2000. Finally, Greenwald is active in promoting diversity in Computer Science, leading multiple K-12 initiatives in which Brown undergraduates teach computer science to Providence public school students.
2:20 PM – 3:00 PMOn the Role of Mechanism Design in Recommender Ecosystems
Craig Boutilier - Google
Recommender systems (RSs) lie at the center of complex ecosystems, involving large numbers of users, content providers or vendors, advertisers and even competing platforms, whose behaviors are driven by their incentives or preferences for RS-induced outcomes. The resulting interactions can generate complex dynamics which, in turn, impacts the ability of the RS to act in the best interests of any particular actor or implement tradeoffs w.r.t. the interests of different actors. The design of RSs in such settings has received relatively scant attention. We briefly illustrate examples of such interactions and discuss the use of mechanism design (MD)—and—adjacent areas, such as preference elicitation, behavioral economics, reinforcement learning, etc.—as a means to ensure RSs have positive societal impact. We also identify a number of research challenges that must be addressed to bring MD to bear on recommender ecosystems.
Speaker Bio: Craig Boutilier is Principal Scientist at Google. He works on various aspects of decision making under uncertainty, with a current focus on recommender systems (and ecosystems), user modeling, reinforcement learning, preference elicitation, and related topics.
He was a Professor in the Department of Computer Science at the University of Toronto (on leave) and Canada Research Chair in Adaptive Decision Making for Intelligent Systems. He received his Ph.D. in Computer Science from the University of Toronto in 1992, and worked as an Assistant and Associate Professor at the University of British Columbia from 1991 until his return to Toronto in 1999. He served as Chair of the Department of Computer Science at Toronto from 2004-2010. He was co-founder (with Tyler Lu) of Granata Decision Systems from 2012-2015, until his move to Google in 2015.
Boutilier was a consulting professor at Stanford University from 1998-2000, an adjunct professor at the University of British Columbia from 1999-2010, and a visiting professor at Brown University in 1998, at the University of Toronto in 1997-98, at Carnegie Mellon University in 2008-09, and at Université Paris-Dauphine (Paris IX) in the spring of 2011. He served on the Technical Advisory Board of CombineNet, Inc. from 2001 to 2010.
Boutilier's research interests have spanned a wide range of topics, from knowledge representation, belief revision, default reasoning, and philosophical logic, to probabilistic reasoning, decision making under uncertainty, multiagent systems, and machine learning. His current research efforts focus on various aspects of decision making under uncertainty: preference elicitation, mechanism design, game theory and multiagent decision processes, economic models, social choice, computational advertising, Markov decision processes, reinforcement learning and probabilistic inference.
Boutilier is a past Editor-in-Chief of the Journal of Artificial Intelligence Research (JAIR). He was a past Associate Editor with the ACM Transactions on Economics and Computation (TEAC), the Journal of Artificial Intelligence Research (JAIR), the Journal of Machine Learning Research (JMLR), and Autonomous Agents and Multiagent Systems (AAMAS); and he has sat on the editorial/advisory boards of several other journals. Boutilier has organized several international conferences and workshops, including his work as Program Chair of the Twenty-first International Joint Conference on Artificial Intelligence (IJCAI-09) and Program Chair of the Sixteenth Conference on Uncertainty in Artificial Intelligence (UAI-2000). He has also served on the conference program committees of roughly 60 leading international conferences.
Boutilier is a Fellow of the Royal Society of Canada (RSC), the Association for Computing Machinery (ACM) and the Association for the Advancement of Artificial Intelligence (AAAI). He was the recipient of the 2018 ACM/SIGAI Autonomous Agents Research Award, He was awarded a Tier Ir Canada Research Chair, an Isaac Walton Killam Research Fellowship, and an IBM Faculty Award. He received the Killam Teaching Award from the University of British Columbia in 1997. He has also received a number of Best Paper awards including: the 2009 IJCAI-JAIR Best Paper Prize (with R. Brafman, C. Domshlak, H. Hoos, D. Poole, from the Journal of Artificial Intelligence Research); the 2014 AIJ Prominent Paper Award (with S. Sanner, from the journal Artificial Intelligence); the 2018 NeurIPS Best Paper Award (w. T. Lu, D. Schuurmans); and the 2022 AIJ Prominent Paper Award (with I. Caragiannis, S. Haber, T. Lu, A. Procaccia and O. Sheffet, from the journal Artificial Intelligence).
3:00 PM – 3:30 PMBreak (30 minutes)
3:30 PM – 4:10 PMIncentivizing Compliance with Algorithmic Instruments
Vasilis Syrgkanis - Stanford University
Randomized experiments can be susceptible to selection bias due to potential non-compliance by the participants. While much of the existing work has studied compliance as a static behavior, we propose a game-theoretic model to study compliance as dynamic behavior that may change over time. In rounds, a social planner interacts with a sequence of heterogeneous agents who arrive with their unobserved private type that determines both their prior preferences across the actions (e.g., control and treatment) and their baseline rewards without taking any treatment. The planner provides each agent with a randomized recommendation that may alter their beliefs and their action selection. We develop a novel recommendation mechanism that views the planner's recommendation as a form of instrumental variable (IV) that only affects an agents' action selection, but not the observed rewards. We construct such IVs by carefully mapping the history -- the interactions between the planner and the previous agents -- to a random recommendation. Even though the initial agents may be completely non-compliant, our mechanism can incentivize compliance over time, thereby enabling the estimation of the treatment effect of each treatment, and minimizing the cumulative regret of the planner whose goal is to identify the optimal treatment.
4:50 PM – 5:30 PMChallenges in Machine Learning and Game Theory for Social Impact (remote presentation)
Fei Fang - Carnegie Mellon University
6:00 PM – 8:00 PMDinner beginning at 6:00 PM
Friday, October 28, 2022
Workshop Talks
8:30 AM – 9:30 AMIntroductions around the Room
9:00 AM – 9:45 AMOverview of work and relevance to IBM
Segev Wasserkrug - IBM Research
9:30 AM – 10:00 AMOverview of the Learning and Games program at the Simons Institute
Vasilis Syrgkanis - Stanford University
9:45 AM – 10:15 AMIdeas for related events at DIMACS
David Pennock - DIMACS
10:00 AM – 10:30 AMBreakfast is available beginning at 8:30
10:15 AM – 10:45 AMBreak (30 minutes)
10:45 AM – 11:30 AMRump Session
At the Rump Session workshop attendees were invited to speak for 5 minutes about their own related research, others' research, an open problem, a new idea, or anything else relevant to the workshop. Use of slides was optional.
List of presenters (with slides when available):
Learning Non-Truthful Mechanisms
Gianluca Brero, Brown UniversityIncentives and Fairness in Societal Resource Allocation
Sanmay Das, George Mason UniversityNear-optimal No-regret Learning Dynamics for Extensive-form Games
Gabriele Farina, Carnegie Mellon UniversityGame Theory for Network Evolution
Jie Gao, Rutgers UniversityAdditive Context-Dependent Preferences
Stephan Jagau, University of California, IrvineSome Methods for Solving Stackelberg Equilibrium in EFGs (TBC)
Chun Kai Ling, Carnegie Mellon UniversityLearning to Learn in Multiagent Reinforcement Learning
Miao Liu, META Reality LabsAdvances in Double Oracle Algorithms
Stephen McAleer, Carnegie Mellon UniversityNegotiation Protocols for Climate Policy
Max Puelma-Touzel, University of MontrealInfluencing Long-Term Behavior in Multiagent Reinforcement Learning
Matthew Riemer, IBM ResearchLearning-augmented Mechanism Design
Xizhi Tan, Drexel UniversityMulti-agent Performative Prediction
Fang-Yi Yu, George Mason UniversityMediation in Extensive-form Games? (TBC)
Brian Zhang, Carnegie Mellon University11:30 AM – 12:00 PMFull group identifies key themes for break out discussions
12:00 PM – 1:00 PMLunch (1 hour)
1:00 PM – 2:15 PMBreakout group discussions
2:15 PM – 2:30 PMBreak (15 minutes)
2:30 PM – 3:15 PMSummaries from Breakout Groups
3:15 PM – 3:30 PMSummary & Next Steps
- Sponsors: IBM Research
-
This will be a two-day workshop.
- Day 1: Lectures and discussion on current relevant state of the art both in academia and in industry
- Day 2: Brainstorming and breakout sessions for discussions and seeding collaborations.
Presentations at the workshop are by invitation and will occur on only the first day of the workshop. Attendance at the workshop is also by invitation, but those who would like to attend may request an invitation from the organizers. The Day 2 breakouts will occur in a smaller space, so we will have greater ability to accommodate requests to attend the Day 1 lectures than to attend both days.
Lectures will be recorded and posted if the presenter grants permission to do so.
