Analyze Text Data with Yellowbrick

Use visual diagnostic tools from Yellowbrick to steer your machine learning workflow

Vectorize text data using TF-IDF

Cluster documents using embedding techniques and appropriate metrics

Welcome to this project-based course on Analyzing Text Data with Yellowbrick. Tasks such as assessing document similarity, topic modelling and other text mining endeavors are predicated on the notion of "closeness" or "similarity" between documents. In this course, we define various distance metrics (e.g. Euclidean, Hamming, Cosine, Manhattan, etc) and understand their merits and shortcomings as they relate to document similarity. We will apply these metrics on documents within a specific corpus and visualize our results. By the end of this course, you will be able to confidently use visual diagnostic tools from Yellowbrick to steer your machine learning workflow, vectorize text data using TF-IDF, and cluster documents using embedding techniques and appropriate metrics.

Data ScienceNatural Language ProcessingMachine LearningPython ProgrammingData Visualization (DataViz)

  1. Introduction and Loading the Corpus

  2. Vectorizing the Documents

  3. Clustering Similar Documents with Squared Euclidean Distance And Euclidean Distance

  4. Manhattan (aka “Taxicab” or “City Block”) Distance

  5. Bray Curtis Dissimilarity and Canberra Distance

  6. Cosine Distance

  7. What Metrics Not to Use

  8. Omitting Class Labels - Using KMeans Clustering

