You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

生成n-gram时报错AttributeError: 'list' object has no attribute 'str'

报错根因

AttributeError: 'list' object has no attribute 'str'由代码里的逻辑混淆直接触发:

  • 报错定位在generate_N_grams函数首行:text.split(" ")是对传入函数的单个Python原生字符串做空格切分,返回值是原生列表类型。.str是Pandas Series专属的字符串访问接口,Python原生列表不存在这个属性,你在split返回的列表后追加.str调用,直接触发属性不存在的错误。
  • 另外你写的数据集处理逻辑存在无效冗余:前两行对news['text_process_1gram']的赋值会被后续操作完全覆盖,没有实际运行意义,可以直接删除。
修正后可运行代码
from nltk.corpus import stopwords
# 提前加载停用词集合,避免每次调用函数重复加载,提升处理效率
stop_words = set(stopwords.words('english'))

def generate_N_grams(text, ngram=1):
  # 删除多余的.str调用,增加空字符串过滤避免生成无意义的ngram
  words = [word for word in text.split(" ") if word not in stop_words and word.strip() != '']  
  temp = zip(*[words[i:] for i in range(0, ngram)])
  ans = [' '.join(ngram) for ngram in temp]
  return ans

# 移除冗余赋值,统一在apply前做类型转换
news['text_process_1gram'] = news['text_process'].astype(str).apply(lambda x: generate_N_grams(x, 1))

如果你原本的处理逻辑是要先按逗号切分文本,可自行调整split的参数,当前修正版本匹配你原函数按空格分词的设计逻辑。

内容的提问来源于stack exchange,提问作者Dumpling

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 17:27:25