原始数据预处理代码单元测试结果不符问题求解
问题修正:文本预处理代码与单元测试匹配问题
问题根源分析
- 正则表达式错误:
remove方法中的re.sub('^[\sA-Za-z0-9]', '', self.fi.readfile())仅匹配并删除文本开头第一个符合[\sA-Za-z0-9]的字符,导致"Sample"被截断为"ample"。 - 缺失模块导入:预处理代码未导入
re模块,运行时会抛出未定义错误。 - 词形还原与大小写不匹配:原代码未对lemma结果做小写转换,且单元测试预期为小写输出;同时
testing的词形还原需确保逻辑正确。 - 单元测试入口错误:测试脚本中
if __name__ == 'main':应为if __name__ == '__main__':,否则测试无法正常执行。
修正后的预处理代码
import pandas as pd import spacy import re # 补充缺失的re模块导入 from spacy.lang.en.stop_words import STOP_WORDS import nltk nlp = spacy.load("en_core_web_md") class fileread: def readfile(self): file_path = 'C:\\Users\\Documents\\Emails\\DEP72303-SYSOUT.txt' with open(file_path, 'r') as text: return text.read() class preprocess: fi = fileread() def remove(self): raw_text = self.fi.readfile() # 修正正则:移除所有非字母、数字、空格的字符(若无需清理特殊字符可直接返回raw_text) return re.sub(r'[^\sA-Za-z0-9]', '', raw_text) def preprocess(self): doc = nlp(self.remove()) filtered = [] for token in doc: if token.is_stop or token.is_punct: continue # 转换为小写,匹配预期输出格式 filtered.append(token.lemma_.lower()) return " ".join(filtered) pre = preprocess() output = pre.preprocess() tokens = output.split("\n") print(tokens)
修正后的单元测试代码
import unittest import os from preprocessing import preprocess class MockFileRead: def readfile(self): with open('test_input.txt', 'r') as text: return text.read() class TestPreprocess(unittest.TestCase): def test_preprocess(self): sample_input = "Sample text content for testing purposes." with open('test_input.txt', 'w') as temp_file: temp_file.write(sample_input) preprocess_instance = preprocess() preprocess_instance.fi = MockFileRead() expected_output = "sample text content test purpose" self.assertEqual(preprocess_instance.preprocess(), expected_output) os.remove('test_input.txt') if __name__ == '__main__': # 修正测试入口 unittest.main()
关键修正说明
- 正则逻辑修复:将
^[\sA-Za-z0-9]改为[^\sA-Za-z0-9](字符集内的^表示取反),实现移除所有非目标字符的需求,而非删除开头第一个字符。 - 小写转换:对lemma结果做小写处理,确保输出与单元测试的小写预期匹配。
- 模块与入口修正:补充
re模块导入,修正测试脚本的入口判断,保证代码可正常运行。
内容的提问来源于stack exchange,提问作者vinamrata
相关产品推荐
相关产品推荐

