Assessing the Quality of Personality Measures Generated by Large Language Models ready for submission · PNAS

Can an LLM provide a low-cost, scalable measure of Big Five traits from student letters, and do these measures meaningfully predict future outcomes? Using pen-pal letters written by ~1,000 seventh-grade students in rural China, we prompt GPT-4o to rate the 60-item Big Five Inventory-2 and treat the LLM as an additional informant alongside students, teachers, and guardians.

AI ratings show high internal consistency, reproduce established demographic correlates, and predict subsequent academic performance and psychological health. In joint multi-informant models, they retain distinct predictive power — with a Shapley–Owen contribution comparable to teacher reports and larger than student or guardian reports. We also use LLM-extracted letter topics and sentiments to interpret the AI's ratings, showing meaningful textual correlates.

The Effects of China–US Tensions on Science: Evidence from the arXiv Dataset work in progress

Do China–US tensions reduce scientific impact, and through which channels? We link ~541,000 arXiv computer-science preprints (2010–2024) to OpenAlex bibliometrics and estimate a continuous difference-in-differences design around the August 2018 NIH "foreign influence" investigation wave. Treatment intensity is the 2017 share of overseas-Chinese diaspora authors in each topic.

High-diaspora topics experienced a persistent citation decline after August 2018, while publication volume did not fall. A placebo using the mainland-China-based author share yields no comparable effect, and triple-difference estimates confirm the decline concentrates among overseas-Chinese-authored papers. First-year citations are unaffected; the decline emerges at 12–23 months and persists, while text-based novelty measures remain stable — consistent with weakened long-run knowledge diffusion rather than a deterioration in research content.