You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python NLTK在Brown语料库中查找have的相关词形

解决Brown语料库中匹配"have"相关词形的问题

Got it, let's tackle this problem! Since the .count() method only works for exact matches, we need a way to capture all inflected forms and contractions of "have" from the Brown Corpus word list. Here are two reliable approaches:

方法一:使用正则表达式匹配明确变体

这种方法直接定义所有你需要匹配的"have"相关词形,简单高效,适合你清楚知道目标变体的场景。

import nltk
from nltk.corpus import brown
import re

# 先下载Brown语料库(首次运行需要)
nltk.download('brown')

# 定义匹配规则:包含所有常见的have变体,忽略大小写
have_pattern = re.compile(r'^(have|has|had|haven\'t|hasn\'t|hadn\'t|having)$', re.IGNORECASE)

# 筛选语料库中符合规则的词
have_related_words = [word for word in brown.words() if have_pattern.match(word)]

# 统计数量并展示结果
total_count = len(have_related_words)
print(f"共找到{total_count}个have相关词形")
print(f"前10个示例:{have_related_words[:10]}")

方法二:用词形还原覆盖所有屈折变化

如果需要更灵活的匹配(比如覆盖罕见的屈折形式),可以用NLTK的词形还原工具,把每个词还原到词根后判断是否为"have"。需要额外处理缩略形式,确保还原准确。

import nltk
from nltk.corpus import brown
from nltk.stem import WordNetLemmatizer
from nltk.corpus import wordnet

# 下载所需的NLTK资源
nltk.download('brown')
nltk.download('wordnet')
nltk.download('averaged_perceptron_tagger')

lemmatizer = WordNetLemmatizer()

def convert_pos_tag(tag):
    # 将NLTK的POS标签转换为WordNet能识别的格式
    if tag.startswith('V'):
        return wordnet.VERB
    elif tag.startswith('N'):
        return wordnet.NOUN
    elif tag.startswith('J'):
        return wordnet.ADJ
    elif tag.startswith('R'):
        return wordnet.ADV
    return wordnet.VERB  # 默认按动词处理

# 处理常见缩略形式的映射
contraction_map = {
    "haven't": "have not",
    "hasn't": "has not",
    "hadn't": "had not",
    "having": "having"
}

have_related_words = []

for word in brown.words():
    # 统一转为小写并处理缩略形式
    lower_word = word.lower()
    cleaned_word = contraction_map.get(lower_word, lower_word)
    
    # 如果是拆分后的短语,取核心动词部分
    if ' ' in cleaned_word:
        cleaned_word = cleaned_word.split()[0]
    
    # 获取POS标签并还原词形
    pos_tag = nltk.pos_tag([cleaned_word])[0][1]
    lemma = lemmatizer.lemmatize(cleaned_word, pos=convert_pos_tag(pos_tag))
    
    if lemma == 'have':
        have_related_words.append(word)

# 统计结果
total_count = len(have_related_words)
print(f"共找到{total_count}个have相关词形")
print(f"前10个示例:{have_related_words[:10]}")

两种方法对比

  • 正则表达式:优点是代码简洁、运行速度快;缺点是需要手动枚举所有要匹配的变体,无法覆盖未定义的罕见形式。
  • 词形还原:优点是能自动识别所有屈折变化的"have"变体;缺点是需要处理缩略形式,代码稍复杂,运行速度略慢。

内容的提问来源于stack exchange,提问作者Michael Baumgarn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:53:10