By clicking “Accept All Cookies,” you agree to the storing of cookies on your device to enhance site navigation and analyze site usage.

Skip to main content

Working Papers

As IDE scholars' projects progress, they share early-stage research through working papers that offer insights into their findings and methodology.

Featured

Working Papers Evaluating for the Long Term: Learnings from Industry

Leif Sigerson (Pinterest), Tom Cunningham (METR), Winston Chou (Netflix), Sana Pandey (MIT CSAIL), Jonathan Stray (UC Berkeley CHAI), Lo-Hua Yuan (Airbnb ), Eytan Bakshy (Meta), Timothy Chan (Statsig), Molly Davies (Pinterest), Maria Dimakopoulou (Uber), Simon Ejdemyr (Netflix), Kenneth Hung (Meta), Nathan Kallus (Netflix & Cornell University), Thu Le (Lyft), A. Demetri Pananos (Datadog), Lee Richardson (Google), Brennan Schaffner (Knight-Georgetown Institute), Rose Tan, Martin Tingley, Nadia Tomova (Booking.com), Panagiotis Toulis (University of Chicago Booth School of Business), Wenjing Zheng (Roblox), Zander Arnao (Knight-Georgetown Institute), Dean Eckles (MIT)

This paper collects and shares industry knowledge on how to make decisions from short-term experiments that are better aligned with long-term outcomes. Based on a daylong workshop with 26 experts from 15 online platforms and 4 universities, it formulates a series of propositions that reflect current industry knowledge.

Working Papers General Social Agents

Useful social science theories predict behavior across settings. However, applying a theory to make predictions in new settings is challenging: rarely can it be done without ad hoc modifications to account for setting-specific factors. We argue that AI agents put in simulations of those novel settings offer an alternative for applying theory, requiring minimal or no modifications.

Working Papers Is there “Secret Sauce” in Large Language Model Development?

Natalia Fischl-Lanzoni

 

Do leading LLM developers possess a proprietary “secret sauce,” or is LLM performance driven by scaling up compute? Using training and benchmark data for 809 models released between 2022 and 2025, the authors estimate scaling-law regressions with release-date and developer fixed effects.