如何检测文本中的英文闭合复合词?是否存在预定义闭合复合词表?
Great question! Detecting closed compound words (like snowball or grandmother) is trickier than open compounds because they’re written as a single token—so dependency parsing (which worked for your open compound extraction) won’t help directly. Let’s break down the solutions:
Yes, there are reliable predefined lists you can use. The most accessible and free option is WordNet, a lexical database of English that includes thousands of closed compounds. It’s integrated with NLTK, making it easy to work with in code.
Other options include paid professional dictionaries like the Oxford Dictionary of English Compounds, but WordNet is more than sufficient for most use cases.
Since closed compounds exist as single tokens, we need to combine dictionary checks with lexical splitting to verify if a word is indeed a compound of two or more valid English words. Here are two practical approaches:
1. Dictionary + Splitting Validation (Basic)
This method first confirms a word is valid English, then tries splitting it at every possible point to see if both parts are also valid words.
First, install NLTK and download WordNet:
pip install nltk
Then run this code:
from nltk.corpus import wordnet import nltk # Download WordNet if you haven't already nltk.download('wordnet') def is_closed_compound(word): # First, confirm the word itself is a valid English word if not wordnet.synsets(word.lower()): return False # Try splitting the word at every possible position (skip trivial splits) for split_idx in range(1, len(word)-1): part1 = word[:split_idx].lower() part2 = word[split_idx:].lower() # Check if both parts are valid words if wordnet.synsets(part1) and wordnet.synsets(part2): return True return False # Example usage sentence = "The snowball rolled past my grandmother and a blackboard" words = sentence.split() closed_compounds = [word for word in words if is_closed_compound(word)] print(closed_compounds) # Output: ['snowball', 'grandmother', 'blackboard']
2. Combine with Part-of-Speech (POS) Tagging (Improved Accuracy)
Most closed compounds are nouns (or sometimes verbs/adjectives). Adding POS filtering reduces false positives (e.g., avoiding words like apple which can’t be split into meaningful parts). We can use spaCy for POS tagging here:
import spacy from nltk.corpus import wordnet import nltk nltk.download('wordnet') nlp = spacy.load("en_core_web_sm") def extract_closed_compounds(sentence): doc = nlp(sentence) closed_compounds = [] for token in doc: # Focus on nouns (adjust POS tags if you need verbs/adjectives too) if token.pos_ not in ["NOUN", "PROPN"]: continue word = token.text.lower() if not wordnet.synsets(word): continue # Check for valid splits for split_idx in range(1, len(word)-1): part1 = word[:split_idx] part2 = word[split_idx:] if wordnet.synsets(part1) and wordnet.synsets(part2): closed_compounds.append(token.text) break # Stop checking splits once we find a valid one return closed_compounds # Example usage sentence = "I made blueberry pancakes and my grandfather fixed the mailbox" print(extract_closed_compounds(sentence)) # Output: ['blueberry', 'pancakes', 'grandfather', 'mailbox']
Notes on Limitations
- Some words might have valid splits but aren’t actually compounds (though this is rare with WordNet’s validation). For edge cases, you might need a more specialized dictionary.
- Hyphenated compounds sometimes appear as closed compounds (e.g.,
longterminstead oflong-term). You can handle this by normalizing hyphens first (e.g., replacing-with empty string) before checking.
内容的提问来源于stack exchange,提问作者Minions

