You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python调用clean_text时出现'module'类型不可迭代错误如何解决

错误原因

你遇到的argument of type 'module' is not iterable错误核心来自以下问题:

  • 你直接导入了contractions模块,没有正确读取它的缩写字典,在执行if word in contractions判断时,右侧是模块对象而非可迭代的字典对象,触发了该报错。
  • 代码本身存在多处语法、逻辑错误,即使解决上述问题也无法正常运行:
    1. 拓展缩写的逻辑中,text = "".join(new_text)被写在了for循环内部,仅会处理第一个单词就拼接覆盖原text,逻辑完全错误
    2. 字符替换行存在语法错误:text- re.sub(r'[_"\-;%()|+&=*%.,!?:#$@\[\]/]',' ', text) 漏写了等于号,应为text = re.sub(...)
    3. 停用词移除后用"".join拼接,会直接把所有单词连在一起没有空格,后续分词逻辑完全失效
    4. 重复调用分词器属于无效操作,词形还原部分嵌套了两层map,会把单词拆分为单个字符进行还原,结果完全不符合预期
    5. remove_stopwords、re、nltk相关依赖你没有作为参数传入函数,也没有在函数内部导入,属于全局变量依赖,运行环境稍有变化就会报错
修复方案

首先确保你正确安装了依赖:

pip install contractions nltk

修改后的完整clean_text函数如下:

import re
import contractions
import nltk
# 首次运行需要下载nltk依赖资源,取消注释下面两行执行一次即可
# nltk.download('stopwords')
# nltk.download('wordnet')
# nltk.download('omw-1.4')
from nltk.corpus import stopwords

def clean_text(text, remove_stopwords_flag=True):
    '''Text Preprocessing '''
    # Convert words to lower case 
    text = text.lower()

    # Expand contractions
    text = text.split()
    new_text= []
    for word in text:
        if word in contractions.contractions_dict:
            new_text.append(contractions.contractions_dict[word])
        else:
            new_text.append(word)
    # 拼接逻辑移到循环外,用空格拼接避免单词连在一起
    text = " ".join(new_text)
  
    # Format words and remove unwanted characters
    text = re.sub(r'https?:\/\/\S+', '', text, flags=re.MULTILINE) 
    text = re.sub(r'\<a href', ' ', text)
    text = re.sub(r'&amp;', '', text)
    # 补全等于号,修正正则匹配逻辑
    text = re.sub(r'[_"\-;%()|+&=*%.,!?:#$@\[\]/]',' ', text)
    text = re.sub(r'<br />', ' ', text)
    text = re.sub(r'\'', ' ', text)
    # 合并多个空格为单个空格
    text = re.sub(r'\s+', ' ', text).strip()

    # remove stopwords
    if remove_stopwords_flag:
        text = text.split()
        stops = set(stopwords.words("english"))
        text = [w for w in text if not w in stops]
        # 用空格拼接
        text = " ".join(text)

    # 仅保留一次有效分词
    text = nltk.WordPunctTokenizer().tokenize(text)

    # Lemmatize each token
    lemm = nltk.stem.WordNetLemmatizer()
    # 去掉嵌套map,直接对每个单词还原
    text = [lemm.lemmatize(word) for word in text]

    return text

调用方式保持不变即可:

sentences_train = list(map(clean_text, sentences_train))

内容的提问来源于stack exchange,提问作者Ramishka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 04:54:00