You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLP新手求助:基于NLTK实现多类别句子抽取的技术方案

Sentence Classification with NLTK: Your Guide to Extracting Specific Categories

Hey there! As an NLP beginner, this task is totally manageable with NLTK—let’s walk through exactly how to tackle it, plus talk about expanding to more subjective categories later.

Step 1: Classifying Your Core Categories with NLTK

You’ve got three main categories to target, and NLTK has straightforward tools for each:

1. 承诺类句子(含will/shall)

For these, combine sentence tokenization and part-of-speech (POS) tagging to reliably identify sentences with those modal verbs. POS tagging helps you avoid false positives (like if "will" is used as a noun, e.g., "the last will and testament").

2. 成本/预算相关句子

These rely on keyword matching paired with tokenization. Define a set of budget-related terms (like budget, cost, expense, fund) and check if any appear in the sentence.

3. 其他类别句子

Simple—this is just all sentences that don’t fall into the first two categories.

Example Code Implementation

Here’s a working snippet that ties this all together:

import nltk
from nltk.tokenize import sent_tokenize, word_tokenize
from nltk.pos_tag import pos_tag

# Download required NLTK resources (run once)
nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')

# Sample input text
sample_text = """We shall deliver the project by end of next month.
The budget for this phase is $50,000, including material costs.
Our team has 5 years of NLP experience.
We will send weekly progress updates.
Unexpected expenses may come up if you change the project scope."""

# Split text into individual sentences
sentences = sent_tokenize(sample_text)

def categorize_sentences(sentences):
    promise = []
    cost_budget = []
    other = []
    
    # Define keywords for cost/budget category
    cost_keywords = {'budget', 'cost', 'expense', 'price', 'fund', 'financial'}
    
    for sent in sentences:
        # Tokenize and tag parts of speech for the sentence
        tokens = word_tokenize(sent)
        tagged_tokens = pos_tag(tokens)
        
        # Check for promise modal verbs (will/shall as modal verbs, POS tag 'MD')
        has_promise = any(word.lower() in {'will', 'shall'} and tag == 'MD' for word, tag in tagged_tokens)
        if has_promise:
            promise.append(sent)
            continue
        
        # Check for cost/budget keywords
        has_cost_term = any(word.lower() in cost_keywords for word in tokens)
        if has_cost_term:
            cost_budget.append(sent)
            continue
        
        # Fallback to other category
        other.append(sent)
    
    return {
        "承诺类": promise,
        "成本预算类": cost_budget,
        "其他类": other
    }

# Run the classification and print results
results = categorize_sentences(sentences)
for category, sents in results.items():
    print(f"\n**{category}:**")
    for sent in sents:
        print(f"- {sent}")

Step 2: Adding Subjective Categories—How Hard Is It?

The difficulty depends on how you define your subjective category:

Low-Difficulty Subjective Categories

If your category has clear rules or keywords (e.g., "customer complaint" sentences containing words like frustrated, dissatisfied, complaint), you can extend the keyword/rule-based approach above with minimal effort.

Higher-Difficulty Subjective Categories

For vague, context-dependent categories (e.g., "positive customer feedback", "risk assessment statements"), you’ll need to use supervised machine learning with NLTK:

  1. Label a dataset: Create a set of sentences tagged with your new subjective category.
  2. Extract features: Use NLTK to turn sentences into features (e.g., word frequencies, POS tags, n-grams).
  3. Train a classifier: NLTK has built-in classifiers like NaiveBayesClassifier or MaxentClassifier that you can train on your labeled data.

Example: Adding a "Positive Sentiment" Category

Using NLTK’s VADER (a pre-trained sentiment analyzer for social media/text), you can quickly add a positive sentiment category:

from nltk.sentiment import SentimentIntensityAnalyzer

nltk.download('vader_lexicon')
sia = SentimentIntensityAnalyzer()

# Update the categorize_sentences function to include positive sentiment
def categorize_sentences_with_subjective(sentences):
    promise = []
    cost_budget = []
    other = []
    positive_sentiment = []
    
    cost_keywords = {'budget', 'cost', 'expense', 'price', 'fund', 'financial'}
    
    for sent in sentences:
        tokens = word_tokenize(sent)
        tagged_tokens = pos_tag(tokens)
        
        has_promise = any(word.lower() in {'will', 'shall'} and tag == 'MD' for word, tag in tagged_tokens)
        if has_promise:
            promise.append(sent)
            continue
        
        has_cost_term = any(word.lower() in cost_keywords for word in tokens)
        if has_cost_term:
            cost_budget.append(sent)
            continue
        
        # Check for positive sentiment
        sentiment_scores = sia.polarity_scores(sent)
        # Compound score > 0.5 indicates strong positive sentiment
        if sentiment_scores['compound'] > 0.5:
            positive_sentiment.append(sent)
        else:
            other.append(sent)
    
    return {
        "承诺类": promise,
        "成本预算类": cost_budget,
        "正面情感类": positive_sentiment,
        "其他类": other
    }

Final Notes

  • Start small: Test your rule-based system first, then iterate as you add more categories.
  • For complex subjective categories, investing time in labeling a quality dataset will make your classifier much more accurate.

内容的提问来源于stack exchange,提问作者Himanshu Parmar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 06:22:54