Can an LLM provide a low-cost, scalable measure of Big Five traits from student letters, and do these measures meaningfully predict future outcomes? Using pen-pal letters written by ~1,000 seventh-grade students in rural China, we prompt GPT-4o to rate the 60-item Big Five Inventory-2 and treat the LLM as an additional informant alongside students, teachers, and guardians.
AI ratings show high internal consistency, reproduce established demographic correlates, and predict subsequent academic performance and psychological health. In joint multi-informant models, they retain distinct predictive power — with a Shapley–Owen contribution comparable to teacher reports and larger than student or guardian reports. We also use LLM-extracted letter topics and sentiments to interpret the AI's ratings, showing meaningful textual correlates.