Elliott Ash

Associate Professor, ETH Zurich  ·  Law | Economics | Data Science

[email protected] (calendly / email norms)

Welcome to my web page. I am a professor at ETH Zurich and Scientific Lead in the Swiss AI Initiative. I am a CEPR Research Affiliate in Political Economy, Associate Editor at Economic Journal, Co-Editor at Journal of Law and Economics, and ETH AI Center Core Faculty. I received an ERC Starting Grant. I am the founder and developer of Tracelaw.

I held previous appointments at New York University (Scholar in Residence), University of Warwick (Assistant Professor), and Princeton University (Postdoc). I earned a Ph.D. in Economics and J.D. from Columbia University, B.A. (Plan II Honors) from University of Texas at Austin, and LL.M. in International Criminal Law from University of Amsterdam.

Topics: Law and Economics, Political Economy  · Methods: Econometrics, Natural Language Processing, Machine Learning

Seminars & Conferences

Teaching Materials

Recent Working Papers

(with Sultan Mehmood and Christoph Goessmann)

A randomized rollout of a custom AI assistant to 1,559 Pakistani trial-court judges: with targeted training, adoption rose and districts resolved about 6 percent more cases per year with no measurable loss of quality.

Abstract

We present the first large-scale field experiment evaluating the integration of generative AI into a national justice system. In partnership with Pakistan’s judiciary, we built JudgeGPT, a custom generative AI assistant designed for Pakistan’s trial courts, and randomized 1,559 judges serving across 118 courts into one of three arms: (i) access to the assistant with targeted training tailored to its use; (ii) access to AI with generic training on technology and law; and (iii) generic training without access to AI. Judges who received AI access together with targeted training on the use of the tool were more likely to adopt it, use it more intensively, and continue using it over time. Their attitudes toward AI also shift: they expect the tool and the targeted training to increase their productivity. Administrative records move in the same direction: districts with greater exposure to treated judges resolve more cases. At median-district exposure, introducing AI with targeted training corresponds to 1,848 additional cases resolved per year, a 6.3 percent increase over the mean. This increase does not appear to come at the expense of reduced quality, as measured by case appeals (which slightly decline) or text-based measures of judicial writing (which slightly improve). Chat-log-based usage patterns suggest that judges use AI mainly to clarify legal concepts and support drafting. Targeted training appears to shift use toward tasks where language models are likely to be more useful, such as text improvement, and away from more open-ended legal queries where responses are more costly to verify. Overall, the results suggest that generative AI can raise public-sector productivity, but that these gains may depend on targeted training that directs use toward tasks for which the tool is better suited.

(with Sergio Galletta, Federico Masera, and Lehan Zhang)

Uses UFC events — where Joe Rogan is lead commentator — as a natural experiment in podcast exposure, showing his post-2021 rightward turn shifted young men's primary voting toward Trump.

Abstract

Online creators have become a central source of political information for young adults in recent years, and over the same period, young men have shifted toward the political right relative to young women. We examine whether these two major developments are connected by focusing on “The Joe Rogan Experience”, the leading long-form podcast, which has a disproportionately male audience. Our identification leverages the timing and location of Ultimate Fighting Championship (UFC) events, for which Rogan has long served as the lead commentator. In a difference-in-differences design, we find that local UFC events trigger a significant increase in Google searches and in time spent consuming Joe Rogan episodes among young men. Applying text analysis to podcast transcripts, we document a pronounced shift in Rogan’s political rhetoric: until 2020, Rogan frequently expressed support for Sanders with left-populist rhetoric; starting in 2021, his commentary became supportive of Trump with a right-populist slant. Using administrative voter files, we show that exposure to UFC events until 2020 increased Democratic voting in primaries with Sanders, while after that, exposure to UFC events increased Republican voting in primaries with Trump with all effects concentrated among young men.

(with Benjamin Kohler, David Zollikofer, Johanna Einsiedler, and Alexander Hoyle)

Tests whether AI agents can reproduce published social-science results from a paper's methods description and data alone, without seeing the code, results, or full paper; agents largely succeed, with failures traceable to both agent errors and underspecified papers.

Abstract

Recent work has used LLM agents to reproduce empirical social science results with access to both the data and code. We broaden this scope by asking: Can they reproduce results given only a paper’s methods description and original data? We develop an agentic reproduction system that extracts structured methods descriptions from papers, runs reimplementations under strict information isolation—agents never see the original code, results, or paper—and enables deterministic, cell-level comparison of reproduced outputs to the original results. An error attribution step traces discrepancies through the system chain to identify root causes. Evaluating four agent scaffolds and four LLMs on 48 papers with human-verified reproducibility, we find that agents can largely recover published results, but performance varies substantially between models, scaffolds, and papers. Root cause analysis reveals that failures stem both from agent errors and from underspecification in the papers themselves.

(with Benjamin Arold, W. Bentley MacLeod, and Suresh Naidu), R&R at Quarterly Journal of Economics

Parses 30,000 Canadian union contracts, 1986–2015, into measurable worker rights and prices them: when wages are taxed more heavily, bargaining shifts toward untaxed rights, worth about 5.7 percent of wages per standard deviation.

Abstract

Collective bargaining agreements (CBAs) specify the contractual rights of unionized workers, but their full legal content has not yet been analyzed by economists. This paper develops novel natural language methods to analyze the empirical determinants and economic value of these rights using a new collection of 30,000 CBAs from Canada in the period 1986-2015. We parse legally binding rights (e.g., “workers shall receive…”) and obligations (e.g., “the employer shall provide…”) from contract text, and validate our measures through evaluation of clause pairs and comparison to firm surveys on HR practices. Using time-varying province-level variation in labor income tax rates, we find that higher taxes increase the share of worker-rights clauses while reducing pre-tax wages in unionized firms, consistent with a substitution effect away from taxed wages toward untaxed rights. Further, an exogenous increase in the value of outside options (from a leave-one-out instrument for labor demand) increases the share of worker rights clauses in CBAs. Combining the regression estimates, we infer that a one-standard-deviation increase in worker rights is valued at about 5.7% of wages.

(with Soumitra Shukla and Jason Sockin)

Half a million Glassdoor interview reports show job candidates read interviews as signals of employer quality: easy interviews lead high-paying candidates to reject offers, and those who accept after an easy interview end up worse matched.

Abstract

Interviews allow employers to learn about workers, but do they also enable workers to learn about firms? Studying 500,000 interview reports from Glassdoor, we find candidates for high-paying jobs are more likely to reject a job offer if they believe the interview was easy. Easy interviews appear to convey poor “fit” as those who accept offers after easy interviews are two-fifths of a standard deviation less satisfied with their jobs and 10 percent less likely to remain with their employer for at least one year. Analysis of interview narratives using large language models reveals difficult interviews signal colleague ability whereas easy interviews convey a nonselective process. In a small-scale randomized field experiment, an exogenous increase in difficulty elevated perceived difficulty and boosted applicant engagement with the vacancy. Interviews offer workers a preview of match quality, highlighting a channel through which labor markets may become less efficient if firms automate hiring with AI.

(with Gloria Gennaro)

The arrival of C-SPAN television cameras in the House in 1979 made members' floor speeches more emotional, and districts with exogenously higher viewership elected more emotive speakers — transparency changed rhetoric more than legislative effort.

Abstract

We study the effect of televised broadcasts of floor debates on the rhetoric and behavior of U.S. Congress Members. First, we show in a differences-in-differences analysis that the introduction of C-SPAN broadcasts in 1979 increased the use of emotional appeals in the House relative to the Senate, where televised floor debates were not introduced until later. Second, we use exogenous variation in C-SPAN channel positioning as an instrument for C-SPAN viewership by Congressional district and show that House Members from districts with exogenously higher C-SPAN viewership are more emotive in floor debates. Looking to electoral pressures as a mechanism, we find the emotionality effect of C-SPAN is strongest in competitive districts. C-SPAN exposure increases the vote share for incumbent Congress Members and citizens’ approval of their job in Congress, and more so among Members who speak more emotionally. Contra accountability models of transparency, C-SPAN has no effect on measures of legislative effort on behalf of constituents, and if anything it reduces a politician’s constituency orientation. We find that local news coverage — that is, mediated rather than direct transparency — has the opposite effect of C-SPAN, increasing legislative effort but with no effect on emotional rhetoric. These results highlight the importance of audience and mediation in the political impacts of higher transparency.

(with Sergio Galletta and Giacomo Opocher), R&R at Economic Journal

A pre-registered experiment around California's 2024 referendums: an AI chatbot with access to official voter guides improved voters' knowledge and engagement and reduced overconfidence, with no clear effect on turnout or vote direction.

Abstract

This study explores the potential for AI-powered chatbots to strengthen democracy by boosting political knowledge and engagement through better access to political information. We develop and evaluate BallotBot, an AI chatbot with access to official voter guide information from the November 2024 referendums in California. In a pre-registered three-wave survey experiment in the weeks around election day, participants (California voters) were randomly assigned to use either BallotBot or a traditional digital voter guide to answer questions about ballot initiatives. BallotBot access improved participants’ ability to answer in-depth questions, reduced overconfidence, lowered the perceived cost of acquiring information for less-informed participants, and fostered greater engagement with political information. However, it had no clear effect on self-reported turnout or the direction of voting.

Selected Publications — Economics

(with Daniel L. Chen and Suresh Naidu), Quarterly Journal of Economics (2026)

The Manne economics course trained nearly half of U.S. federal judges between 1976 and 1999; attendees subsequently wrote more economics-laden opinions, ruled against regulators more often, and imposed harsher criminal sentences.

Abstract

This paper empirically studies the effects of the early law-and-economics movement on the U.S. judiciary. We focus on the Manne Economics Institute for Federal Judges, an intensive economics course that trained almost half of federal judges between 1976 and 1999. Using the universe of published opinions in U.S. Circuit Courts and 1 million District Court criminal sentencing decisions, we estimate the within-judge effect of Manne program attendance. Selection into attendance was limited, as the program was popular among judges of all backgrounds, frequently oversubscribed, and admitted participants on a first-come, first-served basis. We find that after attending economics training, participating judges use more economics language in their opinions, rule against regulatory agencies more often, and impose more severe criminal sentences. We argue that economics, as a rigorous social science, was especially effective in persuading judges.

(with Massimo Morelli and Matia Vannoni), Journal of Political Economy (2025)

Measures U.S. state legislation with text analysis, 1965–2012, and finds a causal growth dividend of legislative output — concentrated in contingent clauses and sectors where incomplete contracts hurt investment most.

Abstract

This paper analyzes the conditions under which more legislation contributes to economic growth. In the context of U.S. states, we apply natural language processing tools to measure legislative flows for the years 1965-2012. We implement a novel shift-share design for text data, where the instrument for legislation is leave-one-out legal-topic flows interacted with pre-treatment legal-topic shares. We find that at the margin, higher legislative output causes more economic growth. Consistent with more complete laws reducing ex-post hold-up, we find that the effect is driven by the use of contingent clauses, is largest in sectors with high relationship-specific investments, and is increasing with local economic uncertainty.

(with Sergio Galletta and Tommaso Giommoni), American Economic Journal: Economic Policy (2025)

Trains gradient-boosted models on audited Brazilian municipal budgets to predict where corruption is likely, giving auditors a targeting tool and researchers a prediction-based corruption measure for unaudited years.

Abstract

Can machine learning support better governance? In the context of Brazilian municipalities, 2001-2012, we have access to detailed accounts of local budgets and audit data on the associated fiscal corruption. Using the budget variables as predictors, we train a tree-based gradient-boosted classifier to predict the presence of corruption in held-out test data. The trained model, when applied to new data, provides a prediction-based measure of corruption which can be used for new empirical analysis or to support policy responses. We validate the empirical usefulness of this measure by replicating, and extending, some previous empirical evidence on corruption issues in Brazil. We then explore how the predictions can be used to support policies toward corruption. Our policy simulations show that, relative to the status quo policy of random audits, a targeted policy guided by the machine predictions could detect more than twice as many corrupt municipalities for the same audit rate.

(with Sam Asher, Aditi Bhowmick, Sandeep Bhupatiraju, Daniel L. Chen, Tanaya Devi, Christoph Goessmann, Paul Novosad, Bilal Siddiqi), Review of Economics and Statistics (2025)

Across 5 million Indian criminal cases with quasi-random judge assignment, and identity detected from names by a neural classifier, estimates tight zero effects of shared gender, religion, or caste proxy on defendant outcomes in the aggregate.

Abstract

We study judicial in-group bias in Indian criminal courts using a newly collected dataset on over 5 million criminal case records from 2010-2018. After detecting gender and religious identity using a neural-net classifier applied to judge and defendant names, we exploit quasi-random assignment of cases to judges to examine whether defendant outcomes are affected by assignment to a judge with a similar identity. In the aggregate, we estimate tight zero effects of in-group bias based on shared gender, religion, and last name (a proxy for caste). We do find limited in-group bias in some (but not all) settings where identity is salient — in particular, we find a small religious in-group bias during Ramadan, and we find shared-name in-group bias when judge and defendant match on a rare last name.

(with Clémentine Abed Meraim, Philine Widmer, and Sergio Galletta), Economic Journal (conditionally accepted)

When Fox News viewership rises exogenously in a market, local newspapers' own coverage shifts toward Fox's slant — not by borrowing its content, but by moving their writing.

Abstract

This paper examines the diffusion of media slant, specifically how partisan content from national cable news affects local newspapers in the U.S., 2005-2008. We use a text-based measure of cable news slant trained on content from Fox News Channel (FNC), CNN, and MSNBC to analyze how local newspapers adopt FNC’s slant over CNN/MSNBC’s. Our findings show that local news becomes more similar to FNC content in response to an exogenous increase in local FNC viewership. This shift is not limited to borrowing from cable news, but rather, local newspapers’ own content changes. Further, cable TV slant polarizes local news content.

(with W. Bentley MacLeod), American Economic Journal: Economic Policy (2024)

U.S. states that introduced mandatory retirement for supreme court judges lowered the bench's average age and raised its output of opinions and the citations they attract.

Abstract

Anecdotal evidence often points to aging as a cause for reduced work performance. This paper provides empirical evidence on this issue in a context where performance is measurable and there is variation in mandatory retirement policies: U.S. state supreme courts. We find that introducing mandatory retirement reduces the average age of working judges and improves court performance, as measured by output (number of published opinions) and legal impact (number of forward citations to those opinions). Consistent with aging effects as a contributing factor, we find that older judges do about the same amount of work as younger judges, but that work is lower-quality as measured by citations. However, the effect of mandatory retirement on performance is much larger than what would be expected from the change in the age distribution, suggesting that the presence of older judges reduces the performance of younger judges.

(with Arianna Ornaghi and Daniel L. Chen), American Economic Journal: Applied Economics (2024)

Measures judges' gender attitudes from stereotyped language in their own opinions; judges scoring higher vote more conservatively in gender-related cases and interact differently with female colleagues.

Abstract

Do gender attitudes influence interactions with female judges in U.S. Circuit Courts? In this paper, we propose a judge-specific measure of gender attitudes based on use of gender-stereotyped language in the judge’s authored opinions. Exploiting quasi-random assignment of judges to cases and conditioning on judges’ characteristics, we validate the measure showing that higher-slant judges vote more conservatively in gender-related cases. Higher-slant judges interact differently with female colleagues: they are more likely to reverse lower-court decisions if the lower-court judge is a woman than a man, are less likely to assign opinions to female judges, and cite fewer female-authored opinions.

(with Michael Poyker), The Economic Journal (2024)

Quasi-random variation in Fox News channel position, applied to nearly 7 million sentencing decisions, shows conservative-news exposure makes judges incarcerate longer — with larger effects for Black defendants.

Abstract

Local exposure to conservative news causes judges to impose harsher criminal sentences. Our evidence comes from an instrumental variables analysis, where randomness in television channel positioning across localities induces exogenous variation in exposure to Fox News Channel. These treatment data on news viewership are taken to outcomes data on almost 7 million criminal sentencing decisions in the United States for the years 2005-2017. Higher Fox News viewership increases incarceration length, and the effect is stronger for black defendants and for drug-related crimes. We can rule out changes in the behavior of police, prosecutors, or potential offenders as significant drivers. Consistent with changes in voter attitudes as the key mechanism, the effect on sentencing harshness is observed for elected (but not appointed) judges. Fox News viewership also increases self-reported beliefs about the importance of drug crime as a social problem. Media Coverage.

(with Sergio Galletta), American Economic Journal: Applied Economics (2023)

Exogenous exposure to Fox News shrinks local-government budgets — both revenues and spending — by improving Republican electoral fortunes, shifting campaign agendas, and moving voters' fiscal preferences directly.

Abstract

This paper shows that partisan cable news broadcasts have a causal effect on the size and composition of budgets in U.S. localities. Using exogenous channel positioning as an instrument for viewership, we show that exposure to the conservative Fox News Channel reduces both revenues and expenditures. Multiple mechanisms drive these results: Fox News improves election chances for local Republicans, alters politician campaign agendas, and directly shifts voter policy preferences on fiscal issues. Consistent with the priorities of small-government conservatism, we find evidence that the reduction in public services is compensated by increased private provision, in particular through higher student attendance at private schools. The “Fox News Effect” is not just limited to vote shares; it also moves policy to the right.

(with Gloria Gennaro) The Economic Journal (2022)

Builds a text-based scale of emotion versus reason and applies it to 6 million congressional speeches, 1858–2014: emotionality spikes in wartime and is highest for patriotism-related topics.

Abstract

We use computational linguistics techniques to study the use of emotion and reason in political discourse. Our new measure of emotionality in language combines lists of emotive and cognitive words, as well as word embeddings, to construct a text-based scale between emotion and reason. After validating the method against human annotations, we apply it to scale 6 million speeches in the U.S. Congressional Record for the years 1858 through 2014. Intuitively, emotionality spikes during times of war and is highest for patriotism-related topics. In the time series, emotionality was relatively low and stable in the earlier years but increased significantly starting in the late 1970s. Comparing Members of Congress to their colleagues, we find that emotionality is higher for Democrats, for women, for ethnic/religious minorities, for members of the opposition party, and for those with relatively extreme policy preferences (either left-wing or right-wing) as measured by roll call votes.

Selected Publications — Political Science

(with Johann Kruemmel and Jonathan B. Slapin), American Journal of Political Science (2024)

Analyzes 544,000 speeches in German state parliaments and finds reactions to them are gendered: men and women receive similar amounts of applause and jeering on average, but the pattern depends on the gender and topic of the speech.

Abstract

Are non-verbal reactions during parliamentary debate gendered? Do male and female Members of Parliament (MPs) experience applause or jeering differently? In short, yes, and the gendered nature of a speech matters. Using an original corpus of over 544,000 speeches given in German state parliaments, we first estimate the gendered nature of parliamentary speeches, then examine how reactions to speeches given by male and female MPs differ. Female and male MPs receive similarly positive and negative reactions to their speeches on average, but they receive different reactions depending on the gendered nature of the speeches. Speeches using language associated with women’s topics receive fewer reactions overall, and even fewer when delivered by men. The gendered nature of parliamentary interjections could affect how women MPs view their position and how women voters view parliament.

(with Germain Gauthier and Philine Widmer), Political Analysis (2023)

An unsupervised pipeline (Relatio) that identifies the actors in a text and the relations between them, quantifying the narratives in millions of documents; applied to economic and political storytelling in Congress.

Abstract

Social scientists have become increasingly interested in how narratives — the stories in fiction, politics, and life — shape beliefs, behavior, and government policies. This paper provides an unsupervised method to quantify latent narrative structures in text documents. Our pipeline identifies coherent entity groups and maps explicit relations between them in the text. We provide an application to the United States Congressional Record to analyze political and economic narratives in recent decades. Our analysis highlights the dynamics, sentiment, polarization, and interconnectedness of narratives in political discourse.

(with Massimo Morelli and Moritz Osnabruegge), Political Analysis (2021)

Trains a topic classifier on labeled texts from one domain (party platforms) and applies it to an unlabeled domain (parliamentary speeches), combining the focus of supervised learning with the coverage of unsupervised methods.

Abstract

We introduce and assess cross-domain topic classification. In this approach, an algorithm learns to classify topics in a labeled source corpus and then extrapolates topics in an unlabeled target corpus from another domain. The advantage over within-domain supervised learning is significant efficiency gains because one can use existing training data. The advantage over unsupervised topic models is that our approach can be more specifically targeted to a research question and that the resulting topics are easier to validate and interpret. We demonstrate the method in the case of labeled party platforms (source corpus) and unlabeled parliamentary speeches (target corpus). Besides the standard within-domain error metrics, we further validate the cross-domain performance by labeling a subset of target-corpus documents. We find that the classifier assigns topics accurately in the parliamentary speeches, although accuracy varies substantially by topic. We also propose a tool for interpreting the topics and diagnosing cross-domain classification. To assess empirical validity, we present two case studies on how electoral rules and parliamentarian gender influence the choice of speech topics.

(with Massimo Morelli and Richard Van Weelden), Journal of Politics (2017)

Models and measures politicians' incentive to 'posture' on divisive issues at the expense of productive ones; U.S. senators up for election spend more floor time on divisive issues.

Abstract

This paper provides a theoretical and empirical analysis of how politicians allocate their time across issues. When voters are uncertain about an incumbent’s preferences, there is a pervasive incentive to “posture” by spending too much time on divisive issues (which are more informative about a politician’s preferences) at the expense of time spent on common-values issues (which provide greater benefit to voters). Higher transparency over the politicians’ choices can exacerbate the distortions. These theoretical results motivate an empirical study of how Members of the U.S. Congress allocate time across issues in their floor speeches. We find that U.S. Senators spend more time on divisive issues when they are up for election, consistent with electorally induced posturing. In addition, we find that U.S. House Members spend more time on divisive issues in response to higher news transparency.

Selected Publications — AI/ML/NLP

(with many co-authors), ACL (2026)

A fully open family of language models trained on 15 trillion tokens across more than 1,800 languages, using only openly licensed data and a training objective that suppresses verbatim memorization.

Abstract

We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today’s open model ecosystem: data compliance and multilingual representation. Unlike many prior models that release weights without reproducible data pipelines or regard for content-owner rights, Apertus models are pretrained exclusively on openly available data, retroactively respecting robots.txt exclusions and filtering for non-permissive, toxic, and personally identifiable content. To mitigate risks of memorization, we adopt the Goldfish objective during pretraining, strongly suppressing verbatim recall of data while retaining downstream task performance. The Apertus models also expand multilingual coverage, training on 15T tokens from over 1800 languages, with ~40% of pretraining data allocated to non-English content. Released at 8B and 70B scales, Apertus approaches state-of-the-art results among fully open models on multilingual benchmarks, rivalling or surpassing open-weight counterparts. Beyond model weights, we release all scientific artifacts from our development cycle with a permissive license, including data preparation scripts, checkpoints, evaluation suites, and training code, enabling transparent audit and extension.

(with Afra Amini, Tim Vieira Ryan Cotterell), ICLR (2025)

Distills Best-of-N sampling — generate N answers, keep the best — into the model itself, cutting the inference cost of alignment by a factor of N while keeping most of the quality.

Abstract

Best-of-N (BoN) is a popular and effective algorithm for aligning language models to human preferences. The algorithm works as follows: at inference time, N samples are drawn from the language model, and the sample with the highest reward, as judged by a reward model, is returned as the output. Despite its effectiveness, BoN is computationally expensive; it reduces sampling throughput by a factor of N. To make BoN more efficient at inference time, one strategy is to fine-tune the language model to mimic what BoN does during inference. To achieve this, we derive the distribution induced by the BoN algorithm. We then propose to fine-tune the language model to minimize backward KL divergence to the BoN distribution. Our approach is analogous to mean-field variational inference and, thus, we term it variational BoN (vBoN). To the extent this fine-tuning is successful and we end up with a good approximation, we have reduced the inference cost by a factor of N. Our experiments on controlled generation and summarization tasks show that BoN is the most effective alignment method, and our variational approximation to BoN achieves the closest performance to BoN and surpasses models fine-tuned using the standard KL-constrained RL objective. In the controlled generation task, vBoN appears more frequently on the Pareto frontier of reward and KL divergence compared to other alignment methods. In the summarization task, vBoN achieves high reward values across various sampling temperatures.

(with Dominik Stammbach, Philine Widmer, Eunjung Cho, Caglar Gulcehre), EMNLP (2024)

Aligns language models to the full range of Swiss party positions using 100,000 candidate statements, producing models that represent diverse political viewpoints instead of a single normative stance, plus a method for balanced multi-party overviews.

Abstract

Large language models such as ChatGPT exhibit striking political biases. If users query them about political information, they often take a normative stance. To overcome this, we align LLMs with diverse political viewpoints from 100,000 comments written by candidates running for national parliament in Switzerland. Models aligned with this data can generate more accurate political viewpoints from Swiss parties, compared to commercial models such as ChatGPT. We also propose a procedure to generate balanced overviews summarizing multiple viewpoints using such models. The replication package contains all code and data.

(with Robert Mahari, Dominik Stammbach, and Alex Pentland), ACL (2024)

Releases millions of examples of U.S. federal judges citing precedent in context, as a benchmark for retrieving relevant precedent passages mid-argument — a task where the best models still miss much of what matters.

Abstract

We present the Legal Passage Retrieval Dataset, LePaRD. LePaRD contains millions of examples of U.S. federal judges citing precedent in context. The dataset aims to facilitate work on legal passage retrieval, a challenging practice-oriented legal retrieval and reasoning task. Legal passage retrieval seeks to predict relevant passages from precedential court decisions given the context of a legal argument. We extensively evaluate various approaches on LePaRD, and find that classification-based retrieval appears to work best. Our best models only achieve a recall of 59% when trained on data corresponding to the 10,000 most-cited passages, underscoring the difficulty of legal passage retrieval. By publishing LePaRD, we provide a large-scale and high quality resource to foster further research on legal passage retrieval. We hope that research on this practice-oriented NLP task will help expand access to justice by reducing the burden associated with legal research via computational assistance. Warning: Extracts from judicial opinions may contain offensive language.

(with Maria Antoniak, Joel Mire, Andrew Piper, and Maarten Sap), ACL (2024)

The StorySeeker toolkit: an annotated dataset of 502 Reddit posts and comments, a codebook, and models that detect where people are telling stories across hundreds of online communities.

Abstract

Story detection in online communities is a challenging task as stories are scattered across communities and interwoven with non-storytelling spans within a single text. We address this challenge by building and releasing the StorySeeker toolkit, including a richly annotated dataset of 502 Reddit posts and comments, a detailed codebook adapted to the social media context, and models to predict storytelling at the document and span levels. Our dataset is sampled from hundreds of popular English-language Reddit communities ranging across 33 topic categories, and it contains fine-grained expert annotations, including binary story labels, story spans, and event spans. We evaluate a range of detection methods using our data, and we identify the distinctive textual features of online storytelling, focusing on storytelling spans, which we introduce as a new task. We illuminate distributional characteristics of storytelling on a large community-centric social media platform, and we also conduct a case study on r/ChangeMyView, where storytelling is used as one of many persuasive strategies, illustrating that our data and models can be used for both inter- and intra-community research. Finally, we discuss implications of our tools and analyses for narratology and the study of online communities.

(with Jingwei Ni, Minjing Shi, Dominik Stammbach, Mrinmaya Sachan, Markus Leippold), ACL (2024)

Uses large language models as annotators for detecting factual claims — the first step of automated fact-checking — with a unified, verifiability-based definition of what counts as a claim.

Abstract

With the rise of generative AI, automated fact-checking methods to combat misinformation are becoming more and more important. However, factual claim detection, the first step in a fact-checking pipeline, suffers from two key issues that limit its scalability and generalizability: (1) inconsistency in definitions of the task and what a claim is, and (2) the high cost of manual annotation. To address (1), we review the definitions in related work and propose a unifying definition of factual claims that focuses on verifiability. To address (2), we introduce AFaCTA (Automatic Factual Claim deTection Annotator), a novel framework that assists in the annotation of factual claims with the help of large language models (LLMs). AFaCTA calibrates its annotation confidence with consistency along three predefined reasoning paths. Extensive evaluation and experiments in the domain of political speech reveal that AFaCTA can efficiently assist experts in annotating factual claims and training high-quality classifiers, and can work with or without expert supervision. Our analyses also result in PoliClaim, a comprehensive claim detection dataset spanning diverse political topics.

(with Dominik Stammbach, Vilém Zouhar, Alexander Hoyle, and Mrinmaya Sachan), EMNLP (2023)

Shows that large language models judge topic-model output closer to how humans do than any existing automatic metric, and can be used to choose the number of topics.

Abstract

Topic models are used to make sense of large text collections. However, automatically evaluating topic model output and determining the optimal number of topics both have been longstanding challenges, with no effective automated solutions to date. This paper proposes using large language models to evaluate such output. We find that large language models appropriately assess the resulting topics, correlating more strongly with human judgments than existing automated metrics. We then investigate whether we can use large language models to automatically determine the optimal number of topics. We automatically assign labels to documents and choosing configurations with the most pure labels returns reasonable values for the optimal number of topics.

(with Florian Dorner, Momchil Peychev, Nikola Konstantinov, Naman Goel, Martin Vechev), ICLR (2023)

Lets humans specify fairness requirements in plain language (e.g., 'gender swaps shouldn't change the outcome'), formalizes them, and trains text classifiers that respect them better than hardcoded word-swapping rules.

Abstract

Text classifiers have promising applications in high-stake tasks such as resume screening and content moderation. These classifiers must be fair and avoid discriminatory decisions by being invariant to perturbations of sensitive attributes such as gender or ethnicity. However, there is a gap between human intuition about these perturbations and the formal similarity specifications capturing them. While existing research has started to address this gap, current methods are based on hardcoded word replacements, resulting in specifications with limited expressivity or ones that fail to fully align with human intuition (e.g., in cases of asymmetric counterfactuals). This work proposes novel methods for bridging this gap by discovering expressive and intuitive individual fairness specifications. We show how to leverage unsupervised style transfer and GPT-3’s zero-shot capabilities to automatically generate expressive candidate pairs of semantically similar sentences that differ along sensitive attributes. We then validate the generated pairs via an extensive crowdsourcing study, which confirms that a lot of these pairs align with human intuition about fairness in the context of toxicity classification. Finally, we show how limited amounts of human feedback can be leveraged to learn a similarity specification that can be used to train downstream fairness-aware models. [ ][ ][ ].

(with Nianlong Gu and Richard Hahnloser), ACL (2022)

An extractive summarizer that picks sentences one at a time while tracking what it has already extracted, matching or beating much larger models on summarizing long documents.

Abstract

We introduce MemSum (Multi-step Episodic Markov decision process extractive SUMmarizer), a reinforcement-learning-based extractive summarizer enriched at any given time step with information on the current extraction history. Similar to previous models in this vein, MemSum iteratively selects sentences into the summary. Our innovation is in considering a broader information set when summarizing that would intuitively also be used by humans in this task: 1) the text content of the sentence, 2) the global text context of the rest of the document, and 3) the extraction history consisting of the set of sentences that have already been extracted. With a lightweight architecture, MemSum nonetheless obtains state-of-the-art test-set performance (ROUGE score) on long document datasets (PubMed, arXiv, and GovReport). Supporting analysis demonstrates that the added awareness of extraction history gives MemSum robustness against redundancy in the source document.

Selected Publications — Law

(2026)

Argues that the two main ways of aligning AI models mirror legal traditions: learning from graded human preferences resembles common-law reasoning from precedent, while learning from rule-based checks resembles civil-law code-following.

Abstract

This chapter explores the links between legal reasoning and the post-training regimes of aligned large language models (LLMs). Reinforcement learning from human feedback (RLHF), which optimizes model outputs to match graded human preference rankings, leads to LLMs that perform value-weighted interpolations across task-response pairs. Reinforcement learning with verifiable rewards (RLVR), which rewards outputs that pass objective external checks, produces LLMs that faithfully follow structured reasoning paths. The RLHF approach parallels the analogical, precedent-based reasoning characteristic of common-law systems, while RLVR mirrors the deterministic code-based reasoning associated with civil law. These analogies inform the design of multi-agent legal AI systems, which should deploy RLVR-based LLMs for compliance with bright-line rules, while deploying RLHF-based systems to apply more subtle standards.

(with Shivam Adarsh, Stefan Bechtold, Barton Beebe and Jeanne Fromer), Journal of Empirical Legal Studies (forthcoming)

Tests whether machine learning can place trademarks on the legal 'spectrum of distinctiveness' (fanciful to generic) the way courts do, automating a central judgment of trademark law.

Abstract

Trademark law protects marks in order to enable firms to signal their products’ qualities to consumers. To qualify for protection, a mark must be able to identify and distinguish goods. U.S. courts typically locate a mark on a “spectrum of distinctiveness” – known as the Abercrombie spectrum – that categorizes marks as fanciful, arbitrary, or suggestive, and thus as “inherently distinctive,” or as descriptive or generic, and thus as not inherently distinctive. This paper explores whether locating trademarks on the Abercrombie spectrum can be automated using current natural-language processing techniques. Using about 1.5 million U.S. trademark registrations between 2012 and 2019 as well as 2.2 million related USPTO office actions, the paper presents a machine-learning model that learns semantic features of trademark applications and predicts whether or not a mark is inherently distinctive. Our model can predict trademark actions with 86% accuracy overall, and it can identify subsets of trademark applications where it is highly certain in its predictions of distinctiveness. We further analyze which features in trademark applications drive our model’s predictions. We then explore the practical and normative implications of our approach. On a practical level, we outline a decision-support system that could, as a “robot trademark clerk,” assist trademark experts in their determination of a trademark’s distinctiveness. Such a system could also help trademark experts understand which features of a trademark application contribute the most towards a trademark’s distinctiveness. On a theoretical level, we discuss the normative limits of the Abercrombie spectrum and propose to move beyond Abercrombie for trademarks whose distinctiveness is uncertain. We discuss how machine-learning projects in the law not only inform us about the aspects of the legal system that may be automated in the future, but also force us to tackle normative tradeoffs that may be invisible otherwise.

(with Christoph Goessmann and Suresh Naidu), Philosophical Transactions of the Royal Society A (2024)

Analyzes the municipal codes of 3,259 U.S. cities and finds legal complexity grows geometrically with population — big cities are disproportionately more codified than small ones.

Abstract

Law sets out the rules for society and the economy, particularly important for interactions between strangers. Legal code is a form of non-rival infrastructure, a public good important for investment and innovation. This paper investigates whether legal code complexity scales with population size in US localities. We analyse a corpus of municipal codes from 3259 cities and measure legal complexity using various metrics, including number of words, bytes, and compressed bytes. We find that legal complexity scales geometrically with jurisdiction population, with a scaling parameter of approximately 0.2 and an 𝑅2R2 of approximately 0.2. The estimated scaling parameter is similar to gross domestic product per capita , consistent with an interpretation of legal codes as regulating social interactions per capita in cities.

(with Aniket Kesari, Suresh Naidu, Lena Song, and Dominik Stammbach), ACM Symposium on Computer Science and Law (2024)

An AI pipeline that generates simplified summaries of judicial opinions; in a survey experiment, readers understood the key features of rulings better from the AI summaries than from traditional expert-written ones.

Abstract

Judicial opinions are written to be persuasive and could build public trust in court decisions, yet they can be difficult for non-experts to understand. We present a pipeline for using an AI assistant to generate simplified summaries of judicial opinions. Compared to existing expert-written summaries, these AI-generated simple summaries are more accessible to the public and more easily understood by non-experts. We show in a survey experiment that the AI summaries help respondents understand the key features of a ruling, and have higher perceived quality, especially for respondents with less formal education.

(with Daniel Walters), Cornell Law Review (2023)

Tests the 'Field of Dreams' case for reviving the nondelegation doctrine — that Congress abdicates because it can, and would legislate more if courts forced it to — against the empirical record.

Abstract

A widely held view for why the Supreme Court would be right to revive the nondelegation doctrine is that Congress has perverse incentives to abdicate its legislative role and evade accountability through the use of delegations, either expressly delineated or implied through statutory imprecision, and that enforcement of the nondelegation doctrine would correct for those incentives. We call this the Field of Dreams Theory—if we build the nondelegation doctrine, Congress will legislate. Unlike originalist arguments for the revival of the nondelegation doctrine, this theory has widespread appeal and is instrumental to the Court’s project of gaining popular acceptance of a greater judicial role in policing congressional decisions regarding delegation. But is it true? In this article, we comprehensively test the theory at the state level, using two original datasets: one comprising all laws passed by state legislatures and the other comprising all nondelegation decisions in the state Supreme Courts. Using a variety of measures and methods, and in contrast with the one existing study on the subject, we do observe at least some statistically measurable decrease in delegation, if only by certain measures. However, when put in context, these findings are underwhelming compared to the predictions of the Field of Dreams Theory. For instance, we observe that, even where it exists, this effect is substantively small and on par with a number of other factors that influence delegation—our best estimate is that nondelegation cases explain about 1.5 percent of the variation in delegation. Moreover, we also find some evidence that is directly contrary to the Field of Dreams Theory—that is, we find evidence that enforcement of the nondelegation doctrine actually leads to more implied delegation in the form of vague and precatory statutory language. These findings have direct relevance to contemporary debate and cases entertaining a revitalization of the nondelegation doctrine in the federal courts. First, the findings that enforcement of the doctrine can prospectively decrease legislative delegation suggest that there may be something to the Field of Dreams Theory, although that in turn raises the stakes of debates over whether less delegation would actually be good for public welfare. Second, even though there is an effect, the weakness of that effect, both in an absolute sense and relative to other factors, undermines the overblown claims that the nondelegation doctrine could fundamentally transform how government works. And finally, our finding that judicial decisions enforcing the nondelegation doctrine can sometimes lead to more implied delegation through imprecise statutory language suggests that there may be unintended consequences from giving the nondelegation doctrine a new lease on life.

(with Jeffrey A. Fagan), Georgetown Law Journal Online (2017)

Traces how data-driven, aggressive minor-crime enforcement regimes took hold not only in big cities but in small American municipalities, with consequences for segregation and inequality there.

Abstract

Modern policing emphasizes advanced statistical metrics, new forms of organizational accountability, and aggressive tactical enforcement of minor crimes as the core of its institutional design. Recent policing research has shown how this policing regime has been woven into the social, political and legal systems in urban areas, but there has been little attention to these policing regimes in smaller areas. In these places, where relationships between citizens, courts and police are more intimate and granular, and local boundaries are closely spaced with considerable flow of persons through spaces, the “new policing” has reached deeply into the everyday lives of predominantly non-white citizens through multiple contacts that lead to an array of legal financial obligations including a wide array of fines and fees. Failure to pay these fees often leads to criminal liability. We examine two faces of modern policing, comparing the Ferguson, Missouri and New York City. We analyze rich and detailed panel data from both places on police stops, citations, warrants, arrests, court dispositions, and penalties, to show the web of social control and legal burdens that these practices create. The data paint a detailed picture of racially discriminatory outcomes at all stages of the process that are common to these two very different social contexts. We link the evidence on the spatial concentration of the racial skew in these policing regimes to patterns of social and spatial segregation, and in turn, to the social, economic and health implications for mobility. We conclude with a discussion of the implications of the “new policing” for constitutional regulation and political reform.

All 78 publications and working papers →

Grad Students, Postdocs & Mentees