🚀 AI Interview Questions with Answers Part 9
81. What is Natural Language Processing NLP, and what are its applications?
Natural Language Processing NLP is a branch of Artificial Intelligence that enables computers to understand, interpret, generate, and respond to human language.
Applications:
• Chatbots and virtual assistants
• Machine translation
• Sentiment analysis
• Text summarization
• Spam detection
• Speech recognition
• Question answering systems
82. What are the main stages of the NLP pipeline?
A typical NLP pipeline consists of the following steps:
1. Text collection
2. Text preprocessing
3. Tokenization
4. Stop-word removal
5. Stemming or Lemmatization
6. Feature extraction e.g., TF-IDF or embeddings
7. Model training
8. Evaluation
9. Deployment
Each stage helps convert raw text into meaningful information for AI models.
83. What is tokenization, and why is it important?
Tokenization is the process of breaking text into smaller units called tokens, such as words, subwords, or characters.
Example:
Sentence: "AI is transforming industries."
Word Tokens: AI, is, transforming, industries
Importance:
• First step in NLP
• Makes text understandable for AI models
• Required for language models like BERT and GPT
84. What is the difference between stemming and lemmatization?
Both techniques reduce words to their base form.
Stemming
• Removes prefixes or suffixes using simple rules
• May produce non-dictionary words
• Faster but less accurate
Examples: Running → Run, Studies → Studi
Lemmatization
• Uses vocabulary and grammar rules
• Produces valid dictionary words
• More accurate but computationally slower
Examples: Running → Run, Studies → Study
85. What are stop words, and why are they removed?
Stop words are commonly used words that usually carry little meaningful information.
Examples: the, is, a, an, in, of, on
Removing stop words:
• Reduces dataset size
• Speeds up processing
• Improves model efficiency in many NLP tasks
Note: They may be retained for tasks where sentence context is important.
86. What is TF-IDF, and how is it calculated?
TF-IDF Term Frequency–Inverse Document Frequency is a statistical technique used to measure how important a word is in a document relative to a collection of documents.
It combines:
• Term Frequency TF: How often a word appears in a document
• Inverse Document Frequency IDF: Reduces the importance of words that appear frequently across many documents
Applications: Search engines, Text classification, Information retrieval, Keyword extraction
87. What is Word2Vec, and how does it generate word embeddings?
Word2Vec is a neural network-based technique that converts words into dense numerical vectors.
Words with similar meanings are placed close together in vector space.
Two architectures: Continuous Bag of Words CBOW, Skip-Gram
Applications: Semantic search, Machine translation, Text classification, Recommendation systems
Post #1832
2.49K
- ❤ 4