
Associate Professor
Harvey Mudd College
Office: McGregor 330
xanda@cs.hmc.edu
pronouns: she/her
I'm an Associate Professor in Computer Science at Harvey Mudd College, a small undergraduate-only college centered on educating STEM students as leaders not only through strong technical knowledge but also deep understanding and appreciation of humanities and the social sciences. My passions are empowering using tools from natural language processing for digital humanities and computational social science work, privacy in text models, and teaching about impact and ethics in all CS classes.
I finished my Ph.D. in Computer Science at Cornell University in 2019. My doctoral work, advised by Prof. David Mimno, focused on understanding and improving latent variable models for analysis of real-world datasets by humanist and social science researchers. This work also included ideas on how to integrate privacy into machine-learning-aided data mining in NLP to protect individuals and ideas.
Prior to starting my Ph.D. at Cornell University, I received my B.S. from Harvey Mudd College in Math and Computer Science. I also worked for a short time on the search team at Yelp.

My research group works on questions of how to make it easier for experts in text collections, particularly those in the humanities and social sciences, to start using text mining tools. Currently, this work specifically focuses on user interface design to train and interact with topic models with a low barrier to entry.
In addition to projects on lowering this barrier to entry, I also am a founding member of Pomona/HMC's EconText Lab, which studies ways to use large-scale text analysis tools to approach economics questions.
I publish under Alexandra Schofield, but prefer Xanda Schofield in person.
* Uncited but relevant to this paper is Raphael Cohen et al.'s 2014 PLoS One paper, Redundancy-Aware Topic Modeling for Patient Record Notes, which develops a topic model, Red-LDA, that aims to combat the effects of text duplication.
** This work focuses on solely English, a fact which the title does not specify. While the analysis methodologies should generalize to other languages, the results discouraging stemming may not. See Chandler May et al.'s 2016 arXiv paper Analysis of Morphology in Topic Modeling for an example of a somewhat different result in Russian.
*** This paper uses a version of gender analysis (assuming binary genders and classifying gender using common baby name lists) that I would not recommend due to its lack of gender inclusiveness and inaccuracy. For a better approach, check out Brian Larson's 2017 EthNLP paper, Gender as a Variable in Natural-Language Processing: Ethical Considerations. I have since updated it to include a disclaimer about this practice.