You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用正则表达式修改TXT文件行的问题及脚本优化需求

优化后的TXT清理脚本

核心优化内容

  • 扩展过滤规则,删除包含CPU关键词的行
  • 修复行首冗余前缀清理问题(替代无效的lstrip方法)

待处理TXT示例

[K --More-- [K2024-05-20 10:00:00 系统启动
[K --More-- [KCPU使用率:75%
[K --More-- [K------------------------
[K --More-- [K用户登录成功

[K --More-- [K内存占用:40%

优化后完整代码

import re

def clean_txt(input_path, output_path):
    # 匹配行首的冗余前缀:[K开头 + --More-- + 任意空白 + [K
    prefix_pattern = re.compile(r'^\[K --More--\s+\[K')
    # 匹配纯连字符行
    hyphen_pattern = re.compile(r'^-+$')
    # 匹配日期行(YYYY-MM-DD HH:MM:SS格式)
    date_pattern = re.compile(r'^\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}')

    with open(input_path, 'r', encoding='utf-8') as f:
        lines = f.readlines()
    
    cleaned_lines = []
    for line in lines:
        # 先移除换行符
        line = line.rstrip('\n')
        # 清理行首冗余前缀
        line = prefix_pattern.sub('', line)
        stripped_line = line.strip()
        
        # 跳过空行、纯连字符行、日期行、含CPU的行
        if not stripped_line:
            continue
        if hyphen_pattern.match(stripped_line):
            continue
        if date_pattern.match(stripped_line):
            continue
        if 'CPU' in stripped_line:
            continue
        
        cleaned_lines.append(stripped_line)
    
    with open(output_path, 'w', encoding='utf-8') as f:
        f.write('\n'.join(cleaned_lines))

if __name__ == '__main__':
    clean_txt('input.txt', 'output.txt')

关键优化说明

  1. 行首前缀修复:
    之前lstrip("[K")无效是因为它仅移除开头的单个[或K字符,无法匹配完整前缀。改用正则^\[K --More--\s+\[K精准匹配行首冗余内容,通过sub方法彻底替换为空。

  2. CPU行过滤:
    直接判断处理后的行是否包含CPU关键词,简单高效;若需整词匹配(避免误匹配类似CPUID的字符串),可改用正则r'\bCPU\b'。

  3. 原有逻辑保留:
    保留了空行、纯连字符行、日期行的过滤逻辑,确保原有功能不受影响。

测试输出结果

用户登录成功
内存占用:40%

内容的提问来源于stack exchange,提问作者Martin Barbieri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 00:10:33