You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于语料库词频为Spacy添加自定义停用词的技术咨询

当然可以自定义spaCy来实现这个需求!下面是几个实用的方案,一步步帮你搞定基于文档频率的停用词筛选和集成:

第一步:计算语料库中词汇的文档频率

首先我们需要统计每个词汇在多少文档中出现,再计算其占总文档数的比例,筛选出超过阈值(比如70%)的词汇作为自定义停用词。用spaCy自带的分词能力会比手动写分词函数更准确,因为它会考虑语境和语言规则:

import spacy
from collections import defaultdict

# 加载spaCy模型(根据你的语言选择,这里用英文示例)
nlp = spacy.load("en_core_web_sm")

# 假设你的领域语料库是一个文档列表
domain_docs = [
    "The patient reported symptoms of fatigue",
    "This patient was admitted for observation",
    # 更多领域文档...
]

# 初始化计数器和总文档数
doc_frequency = defaultdict(int)
total_docs = len(domain_docs)

# 遍历所有文档统计词频
for text in domain_docs:
    doc = nlp(text)
    # 提取当前文档中不重复的有效词汇(过滤标点、默认停用词)
    unique_tokens = set(token.text.lower() for token in doc if token.is_alpha and not token.is_stop)
    for token in unique_tokens:
        doc_frequency[token] += 1

# 筛选出文档占比超过70%的词汇作为自定义停用词
threshold = 0.7
custom_stopwords = {token for token, count in doc_frequency.items() if count / total_docs >= threshold}
第二步:将自定义停用词集成到spaCy

有两种灵活的方式可以把这些停用词加入spaCy的处理流程:

方法一:直接修改spaCy的全局停用词表

这种方式会让这些词汇在后续所有处理中都被标记为停用词(token.is_stop = True),适合全局生效的场景:

# 将自定义停用词添加到spaCy的词汇表中
for word in custom_stopwords:
    nlp.vocab[word].is_stop = True

# 测试效果
test_doc = nlp("The patient visited the clinic")
for token in test_doc:
    print(f"Token: {token.text}, Is Stop: {token.is_stop}")
# 输出里"patient"的Is Stop会是True

方法二:自定义Pipeline组件动态过滤

如果你不想修改全局词汇表,而是希望在处理流程中动态过滤这些停用词,可以自定义一个Pipeline组件,这样更灵活,不会影响其他任务的停用词设置:

@spacy.Language.component("custom_stopword_filter")
def custom_stopword_filter(doc):
    # 保留非默认停用词且不在自定义停用词列表中的token
    filtered_tokens = [
        token for token in doc 
        if not token.is_stop and token.text.lower() not in custom_stopwords
    ]
    # 基于过滤后的token生成新的Doc对象
    return doc.from_array(
        doc.vocab,
        [[token.idx, token.text, token.lemma_, token.pos_, token.tag_, token.dep_, token.head.i, token.children] for token in filtered_tokens]
    )

# 将自定义组件添加到spaCy的Pipeline中(放在分词、标注之后)
nlp.add_pipe("custom_stopword_filter", after="tagger")

# 测试效果
test_doc = nlp("The patient visited the clinic")
print([token.text for token in test_doc])
# 输出会自动去掉"patient"
结合Pandas的实现方式

如果你更习惯用Pandas处理语料,也可以结合apply来实现统计,本质和上面的逻辑一致:

import pandas as pd

# 把语料放入DataFrame
df = pd.DataFrame({"text": domain_docs})

# 用spaCy分词并提取每个文档的唯一有效词汇
df["unique_tokens"] = df["text"].apply(
    lambda x: set(token.text.lower() for token in nlp(x) if token.is_alpha and not token.is_stop)
)

# 统计每个词汇的文档出现次数
doc_freq = df["unique_tokens"].explode().value_counts()
# 筛选出超过阈值的停用词
custom_stopwords = set(doc_freq[doc_freq / total_docs >= threshold].index)

之后再用上面两种方法之一把这些停用词集成到spaCy即可。

内容的提问来源于stack exchange,提问作者J-H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:58:57