基于语料库词频为Spacy添加自定义停用词的技术咨询
当然可以自定义spaCy来实现这个需求!下面是几个实用的方案,一步步帮你搞定基于文档频率的停用词筛选和集成:
第一步:计算语料库中词汇的文档频率
首先我们需要统计每个词汇在多少文档中出现,再计算其占总文档数的比例,筛选出超过阈值(比如70%)的词汇作为自定义停用词。用spaCy自带的分词能力会比手动写分词函数更准确,因为它会考虑语境和语言规则:
import spacy from collections import defaultdict # 加载spaCy模型(根据你的语言选择,这里用英文示例) nlp = spacy.load("en_core_web_sm") # 假设你的领域语料库是一个文档列表 domain_docs = [ "The patient reported symptoms of fatigue", "This patient was admitted for observation", # 更多领域文档... ] # 初始化计数器和总文档数 doc_frequency = defaultdict(int) total_docs = len(domain_docs) # 遍历所有文档统计词频 for text in domain_docs: doc = nlp(text) # 提取当前文档中不重复的有效词汇(过滤标点、默认停用词) unique_tokens = set(token.text.lower() for token in doc if token.is_alpha and not token.is_stop) for token in unique_tokens: doc_frequency[token] += 1 # 筛选出文档占比超过70%的词汇作为自定义停用词 threshold = 0.7 custom_stopwords = {token for token, count in doc_frequency.items() if count / total_docs >= threshold}
第二步:将自定义停用词集成到spaCy
有两种灵活的方式可以把这些停用词加入spaCy的处理流程:
方法一:直接修改spaCy的全局停用词表
这种方式会让这些词汇在后续所有处理中都被标记为停用词(token.is_stop = True),适合全局生效的场景:
# 将自定义停用词添加到spaCy的词汇表中 for word in custom_stopwords: nlp.vocab[word].is_stop = True # 测试效果 test_doc = nlp("The patient visited the clinic") for token in test_doc: print(f"Token: {token.text}, Is Stop: {token.is_stop}") # 输出里"patient"的Is Stop会是True
方法二:自定义Pipeline组件动态过滤
如果你不想修改全局词汇表,而是希望在处理流程中动态过滤这些停用词,可以自定义一个Pipeline组件,这样更灵活,不会影响其他任务的停用词设置:
@spacy.Language.component("custom_stopword_filter") def custom_stopword_filter(doc): # 保留非默认停用词且不在自定义停用词列表中的token filtered_tokens = [ token for token in doc if not token.is_stop and token.text.lower() not in custom_stopwords ] # 基于过滤后的token生成新的Doc对象 return doc.from_array( doc.vocab, [[token.idx, token.text, token.lemma_, token.pos_, token.tag_, token.dep_, token.head.i, token.children] for token in filtered_tokens] ) # 将自定义组件添加到spaCy的Pipeline中(放在分词、标注之后) nlp.add_pipe("custom_stopword_filter", after="tagger") # 测试效果 test_doc = nlp("The patient visited the clinic") print([token.text for token in test_doc]) # 输出会自动去掉"patient"
结合Pandas的实现方式
如果你更习惯用Pandas处理语料,也可以结合apply来实现统计,本质和上面的逻辑一致:
import pandas as pd # 把语料放入DataFrame df = pd.DataFrame({"text": domain_docs}) # 用spaCy分词并提取每个文档的唯一有效词汇 df["unique_tokens"] = df["text"].apply( lambda x: set(token.text.lower() for token in nlp(x) if token.is_alpha and not token.is_stop) ) # 统计每个词汇的文档出现次数 doc_freq = df["unique_tokens"].explode().value_counts() # 筛选出超过阈值的停用词 custom_stopwords = set(doc_freq[doc_freq / total_docs >= threshold].index)
之后再用上面两种方法之一把这些停用词集成到spaCy即可。
内容的提问来源于stack exchange,提问作者J-H
相关产品推荐
相关产品推荐

