You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何检测文本中的英文闭合复合词?是否存在预定义闭合复合词表?

Great question! Detecting closed compound words (like snowball or grandmother) is trickier than open compounds because they’re written as a single token—so dependency parsing (which worked for your open compound extraction) won’t help directly. Let’s break down the solutions:

Predefined Closed Compound Word Lists

Yes, there are reliable predefined lists you can use. The most accessible and free option is WordNet, a lexical database of English that includes thousands of closed compounds. It’s integrated with NLTK, making it easy to work with in code.

Other options include paid professional dictionaries like the Oxford Dictionary of English Compounds, but WordNet is more than sufficient for most use cases.

Methods to Detect Closed Compounds in Text

Since closed compounds exist as single tokens, we need to combine dictionary checks with lexical splitting to verify if a word is indeed a compound of two or more valid English words. Here are two practical approaches:

1. Dictionary + Splitting Validation (Basic)

This method first confirms a word is valid English, then tries splitting it at every possible point to see if both parts are also valid words.

First, install NLTK and download WordNet:

pip install nltk

Then run this code:

from nltk.corpus import wordnet
import nltk

# Download WordNet if you haven't already
nltk.download('wordnet')

def is_closed_compound(word):
    # First, confirm the word itself is a valid English word
    if not wordnet.synsets(word.lower()):
        return False
    # Try splitting the word at every possible position (skip trivial splits)
    for split_idx in range(1, len(word)-1):
        part1 = word[:split_idx].lower()
        part2 = word[split_idx:].lower()
        # Check if both parts are valid words
        if wordnet.synsets(part1) and wordnet.synsets(part2):
            return True
    return False

# Example usage
sentence = "The snowball rolled past my grandmother and a blackboard"
words = sentence.split()
closed_compounds = [word for word in words if is_closed_compound(word)]
print(closed_compounds)  # Output: ['snowball', 'grandmother', 'blackboard']

2. Combine with Part-of-Speech (POS) Tagging (Improved Accuracy)

Most closed compounds are nouns (or sometimes verbs/adjectives). Adding POS filtering reduces false positives (e.g., avoiding words like apple which can’t be split into meaningful parts). We can use spaCy for POS tagging here:

import spacy
from nltk.corpus import wordnet
import nltk

nltk.download('wordnet')
nlp = spacy.load("en_core_web_sm")

def extract_closed_compounds(sentence):
    doc = nlp(sentence)
    closed_compounds = []
    for token in doc:
        # Focus on nouns (adjust POS tags if you need verbs/adjectives too)
        if token.pos_ not in ["NOUN", "PROPN"]:
            continue
        word = token.text.lower()
        if not wordnet.synsets(word):
            continue
        # Check for valid splits
        for split_idx in range(1, len(word)-1):
            part1 = word[:split_idx]
            part2 = word[split_idx:]
            if wordnet.synsets(part1) and wordnet.synsets(part2):
                closed_compounds.append(token.text)
                break  # Stop checking splits once we find a valid one
    return closed_compounds

# Example usage
sentence = "I made blueberry pancakes and my grandfather fixed the mailbox"
print(extract_closed_compounds(sentence))  # Output: ['blueberry', 'pancakes', 'grandfather', 'mailbox']

Notes on Limitations

  • Some words might have valid splits but aren’t actually compounds (though this is rare with WordNet’s validation). For edge cases, you might need a more specialized dictionary.
  • Hyphenated compounds sometimes appear as closed compounds (e.g., longterm instead of long-term). You can handle this by normalizing hyphens first (e.g., replacing - with empty string) before checking.

内容的提问来源于stack exchange,提问作者Minions

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 05:02:31