Corpus preprocessing语料预处理报'expected string or bytes like object'错误如何解决
错误原因
- 输入类型不匹配:
load_sentences返回的是字符串列表(每个元素对应txt文件的一行内容),但preprocessing函数直接将整个列表传入仅支持字符串/字节输入的re.sub()方法,是触发报错的核心原因。 - 逻辑漏洞:代码末尾将空列表
clean_text赋值给tokens,没有将处理后的结果存入返回列表,就算类型问题修复也会返回空值。
修复后的代码
修正preprocessing函数,遍历列表中的每个句子单独处理,同时补全作业要求的基础预处理逻辑:
import nltk from nltk import word_tokenize, PorterStemmer from nltk.corpus import stopwords nltk.download('punkt') nltk.download('stopwords') import re # 提前加载停用词和词干提取器,避免重复加载浪费资源 stop_words = set(stopwords.words('english')) stemmer = PorterStemmer() def preprocessing(corpus): """ 接收句子集合返回清洗后的版本 :return : 存储清洗后句子的字符串列表 :rtype : list(str) """ clean_text = [] for sentence in corpus: # 处理标点和单词的分隔 sentence = re.sub(r"(\w)([.,;:!?'\"”\)])", r"\1 \2", sentence) sentence = re.sub(r"([.,;:!?'\"“\(])(\w)", r"\1 \2", sentence) # 分词 tokens = re.split(r"\s+", sentence) # 过滤非字母字符、统一转小写 tokens = [re.sub(r'[^a-zA-Z]', '', token.lower()) for token in tokens if token.strip()] # 移除停用词 tokens = [token for token in tokens if token not in stop_words] # 词干提取 tokens = [stemmer.stem(token) for token in tokens] # 拼接为清洗后的句子存入结果列表 clean_sentence = ' '.join(tokens) if clean_sentence.strip(): clean_text.append(clean_sentence) return clean_text
调用示例
corpus = load_sentences('1800_sample.txt') clean_corpus = preprocessing(corpus) print(clean_corpus)
内容的提问来源于stack exchange,提问作者Shylo
相关产品推荐
相关产品推荐

