Lab 2: Frequency and N-Grams
Remember in Lab 1 how much work you put into tokenization? Given that tokens are foundational to pretty much all NLP algorithms, wouldn't it be annoying if you had to repeat all that work for every lab? Well, the good news is that tokenization is such a common step in NLP that, unsurprisingly, there's open-source code that will do it for you.
Enter spaCy (don't ask me about the capitalization, that's their "official" branding...): an open-source Python package that implements a number of common NLP algorithms. We'll be relying on spaCy for the rest of the semester to streamline our code and avoid repetitive work. This week, we'll be taking advantage of spaCy's built-in tokenization capabilities.
For some of you, this may be your first time using spaCy—and even if you have used it before, maybe it's been a while and you've forgotten some things. NLP as a field is highly dependent on libraries like spaCy, and being able to navigate your own way around libraries is a key skill we want you to practice. To this end, while we will provide hints and outlines regarding spaCy usage, we won't tell you exactly what to write. If you run into issues or confusion regarding spaCy syntax, you are strongly encouraged to make use of all available resources, including the linked spaCy documentation pages, StackOverflow, search (whether old-fashioned or AI-powered), and of course grutoring and office hours. And don't forget to keep a log of that process in your journal!
There are no TODOs on this page, it's just here to introduce you to spaCy.
But if it's your first time seeing spaCy, you might consider spending a few minutes skimming the spaCy documentation.
If you choose to do so, you can talk about it in your journal!
(When logged in, completion status appears here.)