You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在词干提取数据清洗后获取完整的单词列表

问题根源

你的现有代码存在以下逻辑/语法错误,导致仅输出单行结果:

  1. 直接遍历df["Tag"]这个Pandas Series时,拿到的是每行的完整逗号分隔字符串,并未拆分得到单个单词
  2. 小写转换、特殊符号替换的操作未正确赋值,不生效
  3. 特殊符号替换环节使用了未定义的new_text变量
  4. 词干提取仅作用于循环遍历的最后一个单词,没有覆盖所有词汇
  5. 结果拼接逻辑错误,未正确累加清洗后的单词

修复后可运行代码

from nltk.corpus import stopwords
from nltk.stem import PorterStemmer
import pandas as pd 

def Clean_stop_words(tag_series):
    stop_words = set(stopwords.words('english'))
    stemmer = PorterStemmer()
    all_clean_words = []
    # 遍历每行tag数据,跳过空值
    for line in tag_series.dropna():
        # 拆分单行逗号分隔的单词
        words = [w.strip() for w in line.split(',')]
        for word in words:
            # 转小写
            word_lower = word.lower()
            # 过滤停用词
            if word_lower not in stop_words:
                # 移除特殊符号
                symbols = "!\"#$%&()*+-./:;<=>?@[\]^_`{|}~\n"
                for s in symbols:
                    word_lower = word_lower.replace(s, '')
                # 词干提取
                stem_word = stemmer.stem(word_lower)
                if stem_word.strip():
                    all_clean_words.append(stem_word)
    # 按要求逗号拼接输出
    final_output = ','.join(all_clean_words)
    print(final_output)
    return all_clean_words

# 调用函数
clean_words = Clean_stop_words(df["Tag"])

调整说明

  • 新增每行字符串拆分逻辑,拿到单个单词后再逐个处理
  • 用列表存储所有清洗后的单词,最终统一拼接,避免字符串拼接的逻辑错误和性能问题
  • 修复所有语法错误,确保小写转换、特殊符号清洗、词干提取的每一步都作用于所有单词
  • 新增空值、空字符串过滤逻辑,避免运行报错

内容的提问来源于stack exchange,提问作者d12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 19:54:01