XANDA SCHOFIELD
CV
Meetings
Headshot of Xanda Schofield
Xanda Schofield

Associate Professor
Harvey Mudd College
Office: McGregor 330
xanda@cs.hmc.edu
pronouns: she/her

Teaching (Fall 2026)

  • CS 5 Black - Introduction to Computer Science ● Canvas

News

  • September 2026: I'm looking forward to giving an invited talk at Sonoma State University this fall as part of their Computer Science Colloquium!
  • November 2025: Thanks to Suzanne Rivoire for presenting our work-in-progress paper, "Revealing the Hidden Curriculum: a Multi-Institutional Perspective," at Frontiers in Education.
  • January 2025: Tenure accomplished! As of July 2025, I will be an Associate Professor of Computer Science at Harvey Mudd College.
  • October 2024: Our group's paper, "'My Very Subjective Human Interpretation': Domain Expert Perspectives on Navigating the Text Analysis Loop for Topic Models", was accepted at GROUP 2025. Congrats to many co-authors after a 5-year journey for this work!

About Me

I'm an Associate Professor in Computer Science at Harvey Mudd College, a small undergraduate-only college centered on educating STEM students as leaders not only through strong technical knowledge but also deep understanding and appreciation of humanities and the social sciences. My passions are empowering using tools from natural language processing for digital humanities and computational social science work, privacy in text models, and teaching about impact and ethics in all CS classes.

I finished my Ph.D. in Computer Science at Cornell University in 2019. My doctoral work, advised by Prof. David Mimno, focused on understanding and improving latent variable models for analysis of real-world datasets by humanist and social science researchers. This work also included ideas on how to integrate privacy into machine-learning-aided data mining in NLP to protect individuals and ideas.

Prior to starting my Ph.D. at Cornell University, I received my B.S. from Harvey Mudd College in Math and Computer Science. I also worked for a short time on the search team at Yelp.


Old Blog Posts (from back when I blogged):

  • February 2019: Choose Your Words Wisely (for Topic Models)
  • August 2018: The Talk-First Strategy of Poster Design
  • February 2018: Building Your Network as a Young PhD Student: Emails
  • October 2015: What is Impostor Syndrome?
  • April 2015: Mastering the Software Engineering Interview

Teaching

Please see course pages for my course-specific office hours.

Current Courses (Fall 2026)

  • CS 5 Black - Introduction to Computer Science ● Canvas

Regular Courses

  • CS 5 Gold - Introduction to Computer Science ● last offered Spring 2022
  • CS 123 - Computing Practices, Projects, and People ● last offered Spring 2025
  • CS 140/MATH 168 - Algorithms ● last offered Spring 2025
  • CS 159 - Natural Language Processing ● last offered Spring 2024 (old version from Fall 2021)

Outside Courses

  • TAPI - Text Data Curation
  • Cornell CS 4820 - Introduction to Algorithms
  • Cornell INFO 3350 - Text Mining for History and Literature
  • Outlier.org's Computer Science I in Java
WHISK logo: Workflows for Humanistic Understanding of Statistical Knowledge. Two crossed whisks with an arch over it of the full group name and WHISK written below.

My research group works on questions of how to make it easier for experts in text collections, particularly those in the humanities and social sciences, to start using text mining tools. Currently, this work specifically focuses on user interface design to train and interact with topic models with a low barrier to entry.

In addition to projects on lowering this barrier to entry, I also am a founding member of Pomona/HMC's EconText Lab, which studies ways to use large-scale text analysis tools to approach economics questions.

Past Work


LDA Preprocessing ● Text Analysis Workflows ● Privacy ● Teaching ● Other


LDA Preprocessing

  • Jin Cheevaprawatdomrong, Alexandra Schofield, and Attapol T. Rutherford. More Than Words: Collocation Tokenization for Latent Dirichlet Allocation Models. arXiv, 2021.
  • Alexandra Schofield, Laure Thompson, and David Mimno. Quantifying the effects of text duplication on semantic models. EMNLP, 2017.
  • Alexandra Schofield, Måns Magnusson, Laure Thompson, and David Mimno. Understanding text pre-processing for latent Dirichlet allocation. ACL Workshop for Women in NLP (WiNLP), 2017.
  • Alexandra Schofield, Måns Magnusson, and David Mimno. Pulling out the stops: Rethinking stopword removal for topic models. EACL, 2017.*
  • Alexandra Schofield and David Mimno. Comparing apples to apple: the effects of stemmers on topic models. TACL Vol. 4, 2016.**

Text Analysis Workflows

  • Alexandra Schofield, Siqi Wu, Theo Bayard de Volo, Tatsuki Kuze, Alfredo Gomez, Sharifa Sultana. "My Very Subjective Human Interpretation": Domain Expert Perspectives on Navigating the Text Analysis Loop for Topic Models. GROUP 2025.
  • Alexandra Schofield. "The Possibilities and Limitations of Natural Language Processing for the Humanities." In The Bloomsbury Handbook to the Digital Humanities. 2022.
  • Simon Babb, Dana Harris, Mia Wang, Ingrid Wu, Theo Bayard de Volo, Alfredo Gomez, Tatsuki Kuze, Taeyun Lee, David Mimno, and Alexandra Schofield. Introducing tsLDA: A Workflow-Oriented Topic Modeling Tool. West Coast NLP, 2021.
  • [Poster] [Video]
  • Alicia Eads, Alexandra Schofield, Fauna Mahootian, David Mimno, and Rens Wilderom. Separating the wheat from the chaff: A topic- and keyword-based procedure for identifying research-relevant text. Poetics, 2020. [Prior IC2S2 paper]
  • Theo Bayard de Volo, Alfredo Gomez, Tatsuki Kuze, and Alexandra Schofield. LDA in the Wild: How Practitioners Develop Topic Models. West Coast NLP, 2020. [Poster] [Video]

Privacy

  • Alexandra Schofield, Gregory Yauney, and David Mimno. Combatting the challenges of local privacy for distributional semantics with compression. NeurIPS Workshop for Privacy in Machine Learning (PriML), 2019.
  • Rishi Bommasani, Steven Wu, and Alexandra Schofield. Towards Private Synthetic Text Generation. NeurIPS Workshop for Machine Learning with Guarantees, 2019.
  • Aaron Schein, Zhiwei Steven Wu, Alexandra Schofield, Mingyuan Zhou, and Hanna Wallach. Locally private Bayesian inference for count models. ICML, 2019.
  • Alexandra Schofield, Aaron Schein, Zhiwei Steven Wu, and Hanna Wallach. A variational inference approach for locally private inference of Poisson factorization models. NeurIPS Workshop on Privacy Preserving Machine Learning (PPML), 2018.

Teaching

  • Suzanne Rivoire and Alexandra Schofield. WIP: Revealing the Hidden Curriculum: A Multi-Institutional Perspective. Frontiers in Education Conference (FIE), 2025.
  • Alexandra Schofield, Richard Wicentowski, and Julie Medero. Learning How To Learn NLP: Developing Introductory Concepts Through Scaffolded Discovery. ACL Workshop on Teaching NLP, 2021.
  • Emily M. Bender, Dirk Hovy, and Alexandra Schofield. Integrating Ethics into the NLP Curriculum. ACL Tutorial, 2020. [Slides] [Notes]

Other

  • Nile Phillips, Sathvika Anand, Michelle Lum, Manisha Goel, Michelle Zemel, and Alexandra Schofield. Cheap Talk: Topic Analysis of CSR Themes on Corporate Twitter. Proceedings of the Joint Workshop on Financial NLP and Economics in NLP, 2024.
  • Stephanie Fulcar, Clifford Ashmun, Jeremy Bakken, Anna Ding, JP Walker, Alexandra Schofield, and Sarah Kavassalis. "Knowing Best: Distinguishing Features of Scientific Papers Cited in Clean Air Policy." Poster at AGU Fall Meeting, 2023.
  • Manisha Goel, Alexandra Schofield, and Michelle Zemel. One Shock, Many Disruptions: Firm Experience After India's Demonetization. AAAI Workshop on Knowledge Discovery from Unstructured Data in Financial Services, 2022. (Acknowledgments to EconText students Eloise Burtis, Chris Nardi, and Matthew Ivler for data preparation!)
  • Jack Hessel and Alexandra Schofield. How effective is BERT without word ordering? Implications for language understanding and data privacy. ACL 2021.
  • Alexandra Schofield and Thomas Davidson. Identifying hate speech in social media. ACM Crossroads Magazine (XRDS) Vol 24 (2), 2017.
  • Alexandra Schofield and Leo Mehr. Gender-distinguishing features in film dialogue. NAACL 2016 Workshop on Computational Linguistics for Literature (CLFL), 2016.***
  • Jack Hessel, Alexandra Schofield, Lillian Lee, David Mimno. What do vegans do in their spare time? Latent interest detection in multi-community networks. NeurIPS Networks Workshop, 2015.
  • Robert Keller, Alexandra Schofield, August Toman-Yih, Zachary Merritt, John Elliott. Automating the explanation of jazz chord progressions using idiomatic analysis. Computer Music Journal 37:4, 54-69, 2013.
  • Robert M Keller, August Toman-Yih, Alexandra Schofield, Zachary Merritt. A creative improvisational companion based on idiomatic harmonic bricks. Proc. 3rd ICCC, 2012.


Notes

I publish under Alexandra Schofield, but prefer Xanda Schofield in person.

* Uncited but relevant to this paper is Raphael Cohen et al.'s 2014 PLoS One paper, Redundancy-Aware Topic Modeling for Patient Record Notes, which develops a topic model, Red-LDA, that aims to combat the effects of text duplication.

** This work focuses on solely English, a fact which the title does not specify. While the analysis methodologies should generalize to other languages, the results discouraging stemming may not. See Chandler May et al.'s 2016 arXiv paper Analysis of Morphology in Topic Modeling for an example of a somewhat different result in Russian.

*** This paper uses a version of gender analysis (assuming binary genders and classifying gender using common baby name lists) that I would not recommend due to its lack of gender inclusiveness and inaccuracy. For a better approach, check out Brian Larson's 2017 EthNLP paper, Gender as a Variable in Natural-Language Processing: Ethical Considerations. I have since updated it to include a disclaimer about this practice.