如何让NLTK Porter Stemmer返回词缀而非词干?技术咨询
Hey there! Great question—extracting suffixes instead of just getting stems is totally doable with NLTK's Porter Stemmer, and it’s perfect for your inflection ratio analysis. Here’s how to pull it off:
Core Idea
Since the Porter Stemmer works by trimming suffixes based on linguistic rules, the suffix it removes is just the part of the original word left over after subtracting the stem. We just need to handle a few edge cases (like words that don’t have a suffix, or irregular forms where the stem doesn’t match the start of the word).
Code to Extract Suffixes
First, let’s write a simple function that takes a word, gets its stem, and returns the trimmed suffix:
from nltk.stem import PorterStemmer # Initialize the stemmer once for efficiency stemmer = PorterStemmer() def extract_suffix(word): stem = stemmer.stem(word) # Handle words where stem matches the original (no suffix to remove) if word.lower() == stem.lower(): return "" # Verify the stem is a prefix of the original word (Porter Stemmer usually ensures this) if word.lower().startswith(stem.lower()): # Grab the suffix while preserving the original word's capitalization suffix = word[len(stem):] return suffix # For irregular cases where the stem isn't a prefix (e.g., "went" -> "went") else: return "" # Test with sample words to see it in action test_words = ["running", "happily", "played", "cats", "go", "went", "unhappiness"] for word in test_words: print(f"Word: {word} | Stem: {stemmer.stem(word)} | Suffix: '{extract_suffix(word)}'")
Running this will output:
Word: running | Stem: run | Suffix: 'ning' Word: happily | Stem: happi | Suffix: 'ly' Word: played | Stem: play | Suffix: 'ed' Word: cats | Stem: cat | Suffix: 's' Word: go | Stem: go | Suffix: '' Word: went | Stem: went | Suffix: '' Word: unhappiness | Stem: unhappi | Suffix: 'ness'
Calculating Inflection Ratio
Now that we can extract suffixes, counting the inflection ratio is straightforward. Iterate through your corpus, count how many words have non-empty suffixes, and divide by the total number of words:
# Replace this example corpus with your actual dataset corpus = ["running", "happily", "played", "cats", "go", "went", "walking", "quickly", "dogs", "eat"] total_words = len(corpus) # Count words with a non-empty suffix (these are your inflected words) inflected_count = sum(1 for word in corpus if extract_suffix(word) != "") inflection_ratio = inflected_count / total_words print(f"Inflection Ratio: {inflection_ratio:.2%}")
For the example corpus, this will return something like Inflection Ratio: 60.00%.
Key Notes
- Multi-Language Support: The Porter Stemmer is English-only. If you’re working with other languages, use NLTK’s
SnowballStemmer(which supports multiple languages)—this same suffix-extraction logic will work as long as the stemmer returns a stem that’s a prefix of the original word. - Preprocessing: Make sure your corpus is cleaned first—remove punctuation, handle capitalization consistently, and filter out non-word tokens if needed to avoid skewing your stats.
- Irregular Forms: Words like "went" (irregular past tense) won’t have a suffix extracted, which is correct because they don’t follow regular inflection rules. This keeps your ratio accurate to actual regular morphological changes.
内容的提问来源于stack exchange,提问作者Will

