如何优化Python简易文本预处理程序使其更优雅高效?
优化建议:简化聊天文本预处理代码(含Pandas方案)
先基于常见的新手实现场景,还原你的原代码和测试文本示例:
原代码示例
def preprocess_text(raw_text): processed_lines = [] for line in raw_text.split('\n'): # 去除时间戳 if '[' in line and ']' in line: line = line[line.find(']')+1:] # 去除发送者标识 if ':' in line: line = line[line.find(':')+1:] if line.strip(): processed_lines.append(line.strip()) return '\n'.join(processed_lines) # 测试文本 test_text = """ [2024-05-20 10:15:00] 张三:今天下午有会议吗? [2024-05-20 10:16:00] 李四:有的,3点在三楼会议室 [2024-05-20 10:17:00] 张三:好的,需要准备什么材料? [2024-05-20 10:18:00] 客服001:请携带项目进度报告即可 """ print(preprocess_text(test_text))
核心优化方向
1. 用正则替代字符串切片,提升灵活性
你的代码依赖固定字符位置切割,一旦格式变化(比如时间戳用()包裹、标识用英文冒号:)就会失效。用正则可以一次性匹配并移除时间戳和标识:
import re def preprocess_text_regex(raw_text): # 匹配时间戳+发送者标识,兼容中英文冒号 pattern = r'\[\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}\]\s*[^::]+[::]' # 替换为空字符串 cleaned_text = re.sub(pattern, '', raw_text) # 过滤空行并去除多余空格 return '\n'.join([line.strip() for line in cleaned_text.split('\n') if line.strip()]) # 测试效果 print(preprocess_text_regex(test_text))
2. 用Pandas处理批量场景
如果你的数据量较大(比如从日志文件、Excel读取聊天记录),Pandas能帮你更高效地批量处理、筛选和导出结果:
第一步:将文本转为结构化DataFrame
import pandas as pd # 把测试文本拆分为每行记录,转为DataFrame lines = [line.strip() for line in test_text.split('\n') if line.strip()] df = pd.DataFrame(lines, columns=['raw_message'])
第二步:批量清洗数据
# 提取有效消息列 df['clean_message'] = df['raw_message'].str.replace( r'\[\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}\]\s*[^::]+[::]', '', regex=True ) # 查看清洗后的结果 print(df[['raw_message', 'clean_message']])
第三步:导出结果
如果需要把清洗后的文本保存到文件,Pandas一行就能完成:
# 导出为纯文本 df['clean_message'].to_csv('clean_chat.txt', index=False, header=False) # 或者导出为Excel df.to_excel('clean_chat.xlsx', index=False)
3. 额外优化点
- 把正则模式定义为常量,方便后续修改匹配规则(比如调整时间戳格式)
- 给函数添加参数支持自定义匹配模式,让代码更通用
- 加入异常处理,避免因格式异常导致程序崩溃
内容的提问来源于stack exchange,提问作者Kat
相关产品推荐
相关产品推荐

