You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Pandas DataFrame中移除停用词及指定词汇?

问题解决:移除DataFrame指定列中的停用词与自定义词汇

错误原因分析

1. 初始循环代码的问题

你写的循环里判断条件用的是if j not in b,而b只是自定义词汇列表,根本没用到合并了停用词的c,所以自然不会移除英文停用词。另外这个循环把所有符合条件的元素都放到了一个一维列表里,没保留原DataFrame的每行结构。

2. 列表推导式的错误

  • 分割逻辑错了:原数据是用逗号分割短语,你写的x.split(' , ')是按「空格+逗号+空格」分割,根本匹配不上原数据的格式,导致分割失败。
  • 遍历逻辑错了:for j in i是把每个短语拆成单个字符(比如把"ryzen"拆成r,y,z,e,n),然后判断单个字符是否在停用词列表里——停用词都是完整单词,单个字符肯定不在,但这样会把短语里的单词拆碎,最后拼接出来的就是残缺的字符,这就是你看到"ryzen"变成"rzen"的原因。

正确实现代码

首先确保你已经下载了NLTK的停用词资源,如果没下载,先运行:

import nltk
nltk.download('stopwords')

然后用以下代码处理:

import pandas as pd
from nltk.corpus import stopwords

# 初始化数据
df = pd.DataFrame({'a' : ['ryzen cpu,ryzen 5 5600x,best,amd ryzen,sale',
                         'cpu,ryzen 9 7800x,available,computer for ryzen,new']})

# 加载停用词并合并自定义词汇
stop_words = set(stopwords.words('english'))
custom_words = {'best', 'sale', 'new', 'available'}
all_filter_words = stop_words.union(custom_words)

def filter_text(text):
    # 按逗号分割成单个短语
    phrases = text.split(',')
    processed_phrases = []
    for phrase in phrases:
        # 把短语拆成单词,过滤掉需要移除的词
        words = phrase.strip().split()
        filtered_words = [word for word in words if word.lower() not in all_filter_words]
        # 如果过滤后还有单词,就重新拼接成短语
        if filtered_words:
            processed_phrases.append(' '.join(filtered_words))
    # 把处理后的短语用逗号连接
    return ','.join(processed_phrases)

# 应用处理函数到指定列
df['a'] = df['a'].apply(filter_text)
print(df)

输出结果

a
ryzen cpu,ryzen 5 5600x,amd ryzen
cpu,ryzen 9 7800x,computer ryzen

完全符合预期输出。

内容的提问来源于stack exchange,提问作者Popeye

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 08:23:08