This directory contains tutorials on natural language processing, text analysis, classification, and summarization using IBM Watsonx and NLP frameworks.
Each tutorial in this directory includes its own setup and installation instructions. Please refer to the individual tutorial files for specific requirements.
Common requirements:
- Python 3.10 - 3.13
- IBM watsonx.ai account
- Install dependencies (see Installation above)
- Navigate to this directory:
cd tutorials/09-text-processing-and-nlp - Open and run your first tutorial
Generate concise summaries using LLMs.
- Topics: Summarization techniques, prompt engineering, evaluation
- Time: 30-40 minutes
Traditional and modern summarization approaches.
- Topics: Extractive vs abstractive, NLTK, transformers
- Time: 35-45 minutes
Build text classifiers using PyTorch and transformers.
- Topics: Fine-tuning, BERT, classification, training
- Prerequisites: Optional dependencies (PyTorch)
- Time: 45-60 minutes
Generate unit tests using Watsonx Code Assistant.
- Topics: Test generation, code quality, automation
- Time: 25-35 minutes
- Extractive: Select important sentences from original text
- Abstractive: Generate new sentences capturing key ideas
- Binary: Two classes (spam/not spam)
- Multi-class: Multiple exclusive classes
- Multi-label: Multiple non-exclusive labels
- Text Preprocessing (tokenization, cleaning)
- Feature Extraction (embeddings, TF-IDF)
- Model Training
- Evaluation
- Summarization: News aggregation, document management, email digests
- Classification: Sentiment analysis, topic modeling, spam filtering
- Named Entity Recognition: Extract entities (people, places, organizations)
- Text Generation: Create new content
- Main Repository README
- IBM Watsonx Documentation
- Hugging Face Transformers
- NLTK Documentation
- spaCy Documentation
Found an issue or want to add a new NLP tutorial? See our Contributing Guide.
See the LICENSE file in the repository root.