You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Pandas DataFrame中移除带VERB词性标签的词

问题解决:移除荷兰语文本中的动词

首先你的代码条件逻辑完全搞反了——你写的if not token.is_stop or token.pos_ == 'VERB'表示「只要不是停用词,或者是动词,就保留」,这自然会留下动词。要移除所有带VERB标签的词,同时保留非停用词,正确的条件应该是同时满足「不是停用词」和「不是动词」:

text_article['final'] = text_article['tokens'].apply(lambda text: " ".join(token.lemma_ for token in nlp(text) if not token.is_stop and token.pos_ != 'VERB'))

另外要注意一个潜在问题:如果你的text_article['tokens']是分词后的字符串列表(比如每个单元格是["woord1", "woord2"]这种格式),直接传给nlp()会报错,因为spaCy的nlp管道只接受字符串。这种情况下需要先把列表拼接成空格分隔的字符串:

text_article['final'] = text_article['tokens'].apply(lambda tokens_list: " ".join(token.lemma_ for token in nlp(" ".join(tokens_list)) if not token.is_stop and token.pos_ != 'VERB'))

还有一种更高效的优化方式:既然你已经做了转小写的预处理,不如直接在第一次用spaCy处理原始文本时完成所有步骤(移除数字、标点、停用词、动词),避免重复调用nlp管道(4000+行重复调用会拖慢处理速度):

def preprocess_text(text):
    doc = nlp(text.lower())
    cleaned_tokens = [token.lemma_ for token in doc if not token.is_digit and not token.is_punct and not token.is_stop and token.pos_ != 'VERB']
    return " ".join(cleaned_tokens)

text_article['final'] = text_article['Text'].apply(preprocess_text)

这样既减少了重复处理的开销,也避免了中间步骤可能出现的格式问题。

内容的提问来源于stack exchange,提问作者Annick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 18:06:19