You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLTK文本预处理函数运行报TypeError要求左操作数为字符串而非列表如何解决

错误产生原因

核心错误出现在第一个过滤循环的append操作:

  • 原代码写的是x.append(x),是把列表x自身作为元素追加到x列表里,而非追加符合条件的单词i。这就导致经过第一轮循环后,x列表中存储的全部是列表类型的元素,而非预期的字符串类型的单词。
  • 后续你将x的拷贝赋值给text变量,遍历text时拿到的每个j都是列表类型,用列表类型去判断是否属于停用词(字符串集合)或标点(字符串),自然触发“需要字符串作为左操作数,而非列表”的类型错误。
修复方法
  1. 修正append操作的参数,把x.append(x)改为x.append(i)
  2. 可同时优化冗余逻辑:第一轮已经通过isalnum()过滤了非字母数字的字符,后续不需要再判断是否属于标点符号,减少不必要的运算。
  3. 建议提前加载停用词,避免每次调用函数都重复加载,提升运行性能。

如果仅保留原有代码结构只修改错误点,修改后代码如下:

def text_transform(text):
    text = text.lower()
    text = nltk.word_tokenize(text)
    
    x = []
    for i in text:
        if i.isalnum():
            # 仅修改这里的参数为i即可解决报错
            x.append(i)
            
    text = x[:] 
    x.clear() 
    
    for j in text:
        if j not in stopwords.words('english') and j not in string.punctuation:
            x.append(j)
    return x

优化逻辑后的精简版本参考:

import nltk
import string
from nltk.corpus import stopwords

# 提前加载停用词,避免重复IO开销
stop_words = set(stopwords.words('english'))

def text_transform(text):
    text = text.lower()
    tokens = nltk.word_tokenize(text)
    # 单次循环完成所有过滤逻辑
    return [token for token in tokens if token.isalnum() and token not in stop_words]

内容的提问来源于stack exchange,提问作者Avirup Saha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 15:45:01