Python调用clean_text时出现'module'类型不可迭代错误如何解决
错误原因
你遇到的argument of type 'module' is not iterable错误核心来自以下问题:
- 你直接导入了
contractions模块,没有正确读取它的缩写字典,在执行if word in contractions判断时,右侧是模块对象而非可迭代的字典对象,触发了该报错。 - 代码本身存在多处语法、逻辑错误,即使解决上述问题也无法正常运行:
- 拓展缩写的逻辑中,
text = "".join(new_text)被写在了for循环内部,仅会处理第一个单词就拼接覆盖原text,逻辑完全错误 - 字符替换行存在语法错误:
text- re.sub(r'[_"\-;%()|+&=*%.,!?:#$@\[\]/]',' ', text)漏写了等于号,应为text = re.sub(...) - 停用词移除后用
"".join拼接,会直接把所有单词连在一起没有空格,后续分词逻辑完全失效 - 重复调用分词器属于无效操作,词形还原部分嵌套了两层map,会把单词拆分为单个字符进行还原,结果完全不符合预期
remove_stopwords、re、nltk相关依赖你没有作为参数传入函数,也没有在函数内部导入,属于全局变量依赖,运行环境稍有变化就会报错
- 拓展缩写的逻辑中,
修复方案
首先确保你正确安装了依赖:
pip install contractions nltk
修改后的完整clean_text函数如下:
import re import contractions import nltk # 首次运行需要下载nltk依赖资源,取消注释下面两行执行一次即可 # nltk.download('stopwords') # nltk.download('wordnet') # nltk.download('omw-1.4') from nltk.corpus import stopwords def clean_text(text, remove_stopwords_flag=True): '''Text Preprocessing ''' # Convert words to lower case text = text.lower() # Expand contractions text = text.split() new_text= [] for word in text: if word in contractions.contractions_dict: new_text.append(contractions.contractions_dict[word]) else: new_text.append(word) # 拼接逻辑移到循环外,用空格拼接避免单词连在一起 text = " ".join(new_text) # Format words and remove unwanted characters text = re.sub(r'https?:\/\/\S+', '', text, flags=re.MULTILINE) text = re.sub(r'\<a href', ' ', text) text = re.sub(r'&', '', text) # 补全等于号,修正正则匹配逻辑 text = re.sub(r'[_"\-;%()|+&=*%.,!?:#$@\[\]/]',' ', text) text = re.sub(r'<br />', ' ', text) text = re.sub(r'\'', ' ', text) # 合并多个空格为单个空格 text = re.sub(r'\s+', ' ', text).strip() # remove stopwords if remove_stopwords_flag: text = text.split() stops = set(stopwords.words("english")) text = [w for w in text if not w in stops] # 用空格拼接 text = " ".join(text) # 仅保留一次有效分词 text = nltk.WordPunctTokenizer().tokenize(text) # Lemmatize each token lemm = nltk.stem.WordNetLemmatizer() # 去掉嵌套map,直接对每个单词还原 text = [lemm.lemmatize(word) for word in text] return text
调用方式保持不变即可:
sentences_train = list(map(clean_text, sentences_train))
内容的提问来源于stack exchange,提问作者Ramishka
相关产品推荐
相关产品推荐

