You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Corpus preprocessing语料预处理报'expected string or bytes like object'错误如何解决

错误原因
  • 输入类型不匹配:load_sentences返回的是字符串列表(每个元素对应txt文件的一行内容),但preprocessing函数直接将整个列表传入仅支持字符串/字节输入的re.sub()方法,是触发报错的核心原因。
  • 逻辑漏洞:代码末尾将空列表clean_text赋值给tokens,没有将处理后的结果存入返回列表,就算类型问题修复也会返回空值。
修复后的代码

修正preprocessing函数,遍历列表中的每个句子单独处理,同时补全作业要求的基础预处理逻辑:

import nltk
from nltk import word_tokenize, PorterStemmer
from nltk.corpus import stopwords
nltk.download('punkt')
nltk.download('stopwords')
import re

# 提前加载停用词和词干提取器,避免重复加载浪费资源
stop_words = set(stopwords.words('english'))
stemmer = PorterStemmer()

def preprocessing(corpus):
    """
    接收句子集合返回清洗后的版本
    :return : 存储清洗后句子的字符串列表
    :rtype : list(str)
    """
    clean_text = []
    for sentence in corpus:
        # 处理标点和单词的分隔
        sentence = re.sub(r"(\w)([.,;:!?'\"”\)])", r"\1 \2", sentence)
        sentence = re.sub(r"([.,;:!?'\"“\(])(\w)", r"\1 \2", sentence)
        # 分词
        tokens = re.split(r"\s+", sentence)
        # 过滤非字母字符、统一转小写
        tokens = [re.sub(r'[^a-zA-Z]', '', token.lower()) for token in tokens if token.strip()]
        # 移除停用词
        tokens = [token for token in tokens if token not in stop_words]
        # 词干提取
        tokens = [stemmer.stem(token) for token in tokens]
        # 拼接为清洗后的句子存入结果列表
        clean_sentence = ' '.join(tokens)
        if clean_sentence.strip():
            clean_text.append(clean_sentence)
    return clean_text
调用示例
corpus = load_sentences('1800_sample.txt')
clean_corpus = preprocessing(corpus)
print(clean_corpus)

内容的提问来源于stack exchange,提问作者Shylo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 13:45:03