Get in Touch

Course Outline

Comprehensive training curriculum

  1. Foundations of NLP
    • Core concepts of NLP
    • Major NLP frameworks
    • Commercial use cases for NLP
    • Web scraping for data acquisition
    • Retrieving text data via various APIs
    • Managing text corpora and preserving associated metadata
    • Benefits of Python and an NLTK quick start
  2. Practical Insights into Corpora and Datasets
    • The necessity of corpora
    • Corpus examination
    • Data attribute types
    • File formats for corpora
    • Preparing datasets for NLP tasks
  3. Analyzing Sentence Structure
    • Key NLP components
    • Natural language comprehension
    • Morphological analysis: stems, words, tokens, and speech tags
    • Syntactic analysis
    • Semantic analysis
    • Managing ambiguity
  4. Text Data Preprocessing
    • Raw text corpora
      • Sentence segmentation
      • Stemming raw text
      • Lemmatization of raw text
      • Removal of stop words
    • Raw sentence corpora
      • Word segmentation
      • Word lemmatization
    • Handling Term-Document and Document-Term matrices
    • Tokenizing text into n-grams and sentences
    • Customized preprocessing workflows
  5. Text Data Analysis
    • Essential NLP features
      • Parsers and parsing techniques
      • POS tagging and taggers
      • Named entity recognition
      • N-grams
      • Bag of words
    • Statistical aspects of NLP
      • Linear algebra concepts for NLP
      • Probability theory in NLP
      • TF-IDF
      • Vectorization
      • Encoders and decoders
      • Normalization
      • Probabilistic models
    • Advanced feature engineering in NLP
      • Introduction to word2vec
      • Components of the word2vec model
      • Underlying logic of word2vec
      • Extending the word2vec concept
      • Applying the word2vec model
    • Case study: Automatic text summarization using the Bag of Words approach with simplified and exact Luhn's algorithms
  6. Document Clustering, Classification, and Topic Modeling
    • Document clustering and pattern mining (e.g., hierarchical clustering, k-means)
    • Document comparison and classification using TFIDF, Jaccard, and cosine similarity
    • Document classification via Naïve Bayes and Maximum Entropy
  7. Extracting Significant Text Elements
    • Dimensionality reduction: PCA, SVD, and non-negative matrix factorization
    • Topic modeling and retrieval via Latent Semantic Analysis
  8. Entity Extraction, Sentiment Analysis, and Advanced Topic Modeling
    • Distinguishing positive vs. negative sentiment
    • Item Response Theory
    • Applying POS tagging to identify people, places, and organizations
    • Advanced topic modeling: Latent Dirichlet Allocation
  9. Case Studies
    • Analyzing unstructured user reviews
    • Sentiment classification and visualization of product reviews
    • Identifying usage patterns from search logs
    • Text classification
    • Topic modelling

Requirements

Familiarity with NLP fundamentals and an understanding of how AI is applied in business contexts

 21 Hours

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories