NLP新手求助:基于NLTK实现多类别句子抽取的技术方案
Hey there! As an NLP beginner, this task is totally manageable with NLTK—let’s walk through exactly how to tackle it, plus talk about expanding to more subjective categories later.
Step 1: Classifying Your Core Categories with NLTK
You’ve got three main categories to target, and NLTK has straightforward tools for each:
1. 承诺类句子(含will/shall)
For these, combine sentence tokenization and part-of-speech (POS) tagging to reliably identify sentences with those modal verbs. POS tagging helps you avoid false positives (like if "will" is used as a noun, e.g., "the last will and testament").
2. 成本/预算相关句子
These rely on keyword matching paired with tokenization. Define a set of budget-related terms (like budget, cost, expense, fund) and check if any appear in the sentence.
3. 其他类别句子
Simple—this is just all sentences that don’t fall into the first two categories.
Example Code Implementation
Here’s a working snippet that ties this all together:
import nltk from nltk.tokenize import sent_tokenize, word_tokenize from nltk.pos_tag import pos_tag # Download required NLTK resources (run once) nltk.download('punkt') nltk.download('averaged_perceptron_tagger') # Sample input text sample_text = """We shall deliver the project by end of next month. The budget for this phase is $50,000, including material costs. Our team has 5 years of NLP experience. We will send weekly progress updates. Unexpected expenses may come up if you change the project scope.""" # Split text into individual sentences sentences = sent_tokenize(sample_text) def categorize_sentences(sentences): promise = [] cost_budget = [] other = [] # Define keywords for cost/budget category cost_keywords = {'budget', 'cost', 'expense', 'price', 'fund', 'financial'} for sent in sentences: # Tokenize and tag parts of speech for the sentence tokens = word_tokenize(sent) tagged_tokens = pos_tag(tokens) # Check for promise modal verbs (will/shall as modal verbs, POS tag 'MD') has_promise = any(word.lower() in {'will', 'shall'} and tag == 'MD' for word, tag in tagged_tokens) if has_promise: promise.append(sent) continue # Check for cost/budget keywords has_cost_term = any(word.lower() in cost_keywords for word in tokens) if has_cost_term: cost_budget.append(sent) continue # Fallback to other category other.append(sent) return { "承诺类": promise, "成本预算类": cost_budget, "其他类": other } # Run the classification and print results results = categorize_sentences(sentences) for category, sents in results.items(): print(f"\n**{category}:**") for sent in sents: print(f"- {sent}")
Step 2: Adding Subjective Categories—How Hard Is It?
The difficulty depends on how you define your subjective category:
Low-Difficulty Subjective Categories
If your category has clear rules or keywords (e.g., "customer complaint" sentences containing words like frustrated, dissatisfied, complaint), you can extend the keyword/rule-based approach above with minimal effort.
Higher-Difficulty Subjective Categories
For vague, context-dependent categories (e.g., "positive customer feedback", "risk assessment statements"), you’ll need to use supervised machine learning with NLTK:
- Label a dataset: Create a set of sentences tagged with your new subjective category.
- Extract features: Use NLTK to turn sentences into features (e.g., word frequencies, POS tags, n-grams).
- Train a classifier: NLTK has built-in classifiers like
NaiveBayesClassifierorMaxentClassifierthat you can train on your labeled data.
Example: Adding a "Positive Sentiment" Category
Using NLTK’s VADER (a pre-trained sentiment analyzer for social media/text), you can quickly add a positive sentiment category:
from nltk.sentiment import SentimentIntensityAnalyzer nltk.download('vader_lexicon') sia = SentimentIntensityAnalyzer() # Update the categorize_sentences function to include positive sentiment def categorize_sentences_with_subjective(sentences): promise = [] cost_budget = [] other = [] positive_sentiment = [] cost_keywords = {'budget', 'cost', 'expense', 'price', 'fund', 'financial'} for sent in sentences: tokens = word_tokenize(sent) tagged_tokens = pos_tag(tokens) has_promise = any(word.lower() in {'will', 'shall'} and tag == 'MD' for word, tag in tagged_tokens) if has_promise: promise.append(sent) continue has_cost_term = any(word.lower() in cost_keywords for word in tokens) if has_cost_term: cost_budget.append(sent) continue # Check for positive sentiment sentiment_scores = sia.polarity_scores(sent) # Compound score > 0.5 indicates strong positive sentiment if sentiment_scores['compound'] > 0.5: positive_sentiment.append(sent) else: other.append(sent) return { "承诺类": promise, "成本预算类": cost_budget, "正面情感类": positive_sentiment, "其他类": other }
Final Notes
- Start small: Test your rule-based system first, then iterate as you add more categories.
- For complex subjective categories, investing time in labeling a quality dataset will make your classifier much more accurate.
内容的提问来源于stack exchange,提问作者Himanshu Parmar

