You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

原始数据预处理代码单元测试结果不符问题求解

问题修正:文本预处理代码与单元测试匹配问题

问题根源分析

  1. 正则表达式错误:remove方法中的re.sub('^[\sA-Za-z0-9]', '', self.fi.readfile())仅匹配并删除文本开头第一个符合[\sA-Za-z0-9]的字符,导致"Sample"被截断为"ample"。
  2. 缺失模块导入:预处理代码未导入re模块,运行时会抛出未定义错误。
  3. 词形还原与大小写不匹配:原代码未对lemma结果做小写转换,且单元测试预期为小写输出;同时testing的词形还原需确保逻辑正确。
  4. 单元测试入口错误:测试脚本中if __name__ == 'main':应为if __name__ == '__main__':,否则测试无法正常执行。

修正后的预处理代码

import pandas as pd
import spacy
import re  # 补充缺失的re模块导入
from spacy.lang.en.stop_words import STOP_WORDS
import nltk

nlp = spacy.load("en_core_web_md")

class fileread:
    def readfile(self):
        file_path = 'C:\\Users\\Documents\\Emails\\DEP72303-SYSOUT.txt'
        with open(file_path, 'r') as text:
            return text.read()

class preprocess:
    fi = fileread()
    
    def remove(self):
        raw_text = self.fi.readfile()
        # 修正正则:移除所有非字母、数字、空格的字符(若无需清理特殊字符可直接返回raw_text)
        return re.sub(r'[^\sA-Za-z0-9]', '', raw_text)
    
    def preprocess(self):
        doc = nlp(self.remove())
        filtered = []
        for token in doc:
            if token.is_stop or token.is_punct:
                continue
            # 转换为小写,匹配预期输出格式
            filtered.append(token.lemma_.lower())
        
        return " ".join(filtered)

pre = preprocess()
output = pre.preprocess()
tokens = output.split("\n")
print(tokens)

修正后的单元测试代码

import unittest
import os
from preprocessing import preprocess

class MockFileRead:
    def readfile(self):
        with open('test_input.txt', 'r') as text:
            return text.read()

class TestPreprocess(unittest.TestCase):
    def test_preprocess(self):
        sample_input = "Sample text content for testing purposes."
        with open('test_input.txt', 'w') as temp_file:
            temp_file.write(sample_input)
        
        preprocess_instance = preprocess()
        preprocess_instance.fi = MockFileRead()
        
        expected_output = "sample text content test purpose"
        self.assertEqual(preprocess_instance.preprocess(), expected_output)
        
        os.remove('test_input.txt')

if __name__ == '__main__':  # 修正测试入口
    unittest.main()

关键修正说明

  • 正则逻辑修复:将^[\sA-Za-z0-9]改为[^\sA-Za-z0-9](字符集内的^表示取反),实现移除所有非目标字符的需求,而非删除开头第一个字符。
  • 小写转换:对lemma结果做小写处理,确保输出与单元测试的小写预期匹配。
  • 模块与入口修正:补充re模块导入,修正测试脚本的入口判断,保证代码可正常运行。

内容的提问来源于stack exchange,提问作者vinamrata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 04:25:07