基于Python NLTK在Brown语料库中查找have的相关词形
解决Brown语料库中匹配"have"相关词形的问题
Got it, let's tackle this problem! Since the .count() method only works for exact matches, we need a way to capture all inflected forms and contractions of "have" from the Brown Corpus word list. Here are two reliable approaches:
方法一:使用正则表达式匹配明确变体
这种方法直接定义所有你需要匹配的"have"相关词形,简单高效,适合你清楚知道目标变体的场景。
import nltk from nltk.corpus import brown import re # 先下载Brown语料库(首次运行需要) nltk.download('brown') # 定义匹配规则:包含所有常见的have变体,忽略大小写 have_pattern = re.compile(r'^(have|has|had|haven\'t|hasn\'t|hadn\'t|having)$', re.IGNORECASE) # 筛选语料库中符合规则的词 have_related_words = [word for word in brown.words() if have_pattern.match(word)] # 统计数量并展示结果 total_count = len(have_related_words) print(f"共找到{total_count}个have相关词形") print(f"前10个示例:{have_related_words[:10]}")
方法二:用词形还原覆盖所有屈折变化
如果需要更灵活的匹配(比如覆盖罕见的屈折形式),可以用NLTK的词形还原工具,把每个词还原到词根后判断是否为"have"。需要额外处理缩略形式,确保还原准确。
import nltk from nltk.corpus import brown from nltk.stem import WordNetLemmatizer from nltk.corpus import wordnet # 下载所需的NLTK资源 nltk.download('brown') nltk.download('wordnet') nltk.download('averaged_perceptron_tagger') lemmatizer = WordNetLemmatizer() def convert_pos_tag(tag): # 将NLTK的POS标签转换为WordNet能识别的格式 if tag.startswith('V'): return wordnet.VERB elif tag.startswith('N'): return wordnet.NOUN elif tag.startswith('J'): return wordnet.ADJ elif tag.startswith('R'): return wordnet.ADV return wordnet.VERB # 默认按动词处理 # 处理常见缩略形式的映射 contraction_map = { "haven't": "have not", "hasn't": "has not", "hadn't": "had not", "having": "having" } have_related_words = [] for word in brown.words(): # 统一转为小写并处理缩略形式 lower_word = word.lower() cleaned_word = contraction_map.get(lower_word, lower_word) # 如果是拆分后的短语,取核心动词部分 if ' ' in cleaned_word: cleaned_word = cleaned_word.split()[0] # 获取POS标签并还原词形 pos_tag = nltk.pos_tag([cleaned_word])[0][1] lemma = lemmatizer.lemmatize(cleaned_word, pos=convert_pos_tag(pos_tag)) if lemma == 'have': have_related_words.append(word) # 统计结果 total_count = len(have_related_words) print(f"共找到{total_count}个have相关词形") print(f"前10个示例:{have_related_words[:10]}")
两种方法对比
- 正则表达式:优点是代码简洁、运行速度快;缺点是需要手动枚举所有要匹配的变体,无法覆盖未定义的罕见形式。
- 词形还原:优点是能自动识别所有屈折变化的"have"变体;缺点是需要处理缩略形式,代码稍复杂,运行速度略慢。
内容的提问来源于stack exchange,提问作者Michael Baumgarn
相关产品推荐
相关产品推荐

