You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量处理文本文件:基于total列最大值修正week1/week2列值

批量修正文本文件中样本数据的最大值限制问题

我手上有上百个结构一致的文本文件要处理,每行是多列数据,规则如下:

  • 第三列是total的行,是对应样本的最大值参照行,第五列的值就是该样本的上限
  • 同一样本的week1、week2行,第五列的值不能超过这个上限
  • 现在部分week1、week2的第五列值比上限大1,需要把这些值改成对应的上限值

输入示例

sample1 100A    total   1   1000
sample2 100A    total   1   5584
sample3 100A    total   1   8125
sample4 100A    total   1   59

sample1 .   year    1   1000
sample1 .   week1   20  1001
sample1 .   week2   50  1001

sample2 .   year    1   5584
sample2 .   week1   20  5585
sample2 .   week2   100 5585

sample3 .   year    1   8125
sample3 .   week1   55  8126
sample3 .   week2   100 8126

sample4 .   year    1   59
sample4 .   week1   10  59
sample4 .   week2   8   59

期望输出示例

sample1 100A    total   1   1000
sample2 100A    total   1   5584
sample3 100A    total   1   8125
sample4 100A    total   1   59
                
sample1 .   year    1   1000
sample1 .   week1   20  1000
sample1 .   week2   50  1000
                
sample2 .   year    1   5584
sample2 .   week1   20  5584
sample2 .   week2   100 5584
                
sample3 .   year    1   8125
sample3 .   week1   55  8125
sample3 .   week2   100 8125
                
sample4 .   year    1   59
sample4 .   week1   10  59
sample4 .   week2   8   59

适合初学者的处理方法(Python脚本)

作为编程初学者,用Python处理最直观,代码逻辑清晰,容易上手。

操作步骤

  1. 把所有要处理的文本文件放到同一个文件夹,比如命名为data_files
  2. 务必先备份所有原文件,避免处理失误导致数据丢失
  3. 运行下面的Python脚本,自动批量处理所有文件

代码实现

import os

# 替换成你存放文件的实际路径,Windows系统路径用双反斜杠,比如"C:\\data_files"
folder_path = "./data_files"

# 遍历文件夹内所有文件
for filename in os.listdir(folder_path):
    # 只处理txt后缀的文本文件,根据你的文件后缀调整
    if filename.endswith(".txt"):
        file_path = os.path.join(folder_path, filename)
        print(f"正在处理:{filename}")
        
        # 第一步:读取文件,记录每个样本的total最大值
        sample_max_dict = {}
        with open(file_path, 'r', encoding='utf-8') as f:
            all_lines = f.readlines()
            for line in all_lines:
                # 跳过空行
                if not line.strip():
                    continue
                # 按空格分割行数据(自动处理多个连续空格)
                line_parts = line.strip().split()
                # 第三列是total时,记录样本名和对应最大值
                if len(line_parts) >= 5 and line_parts[2] == 'total':
                    sample_name = line_parts[0]
                    max_val = int(line_parts[4])
                    sample_max_dict[sample_name] = max_val
        
        # 第二步:修改week1/week2行的第五列值
        modified_lines = []
        for line in all_lines:
            # 空行直接保留
            if not line.strip():
                modified_lines.append(line)
                continue
            line_parts = line.strip().split()
            if len(line_parts) >=5:
                sample_name = line_parts[0]
                # 判断是否是week1或week2行
                if line_parts[2] in ['week1', 'week2']:
                    current_val = int(line_parts[4])
                    # 如果当前值超过最大值,替换为最大值
                    if sample_name in sample_max_dict and current_val > sample_max_dict[sample_name]:
                        line_parts[4] = str(sample_max_dict[sample_name])
                        # 重新拼接成行(用四个空格分隔,匹配原文件格式)
                        modified_line = '    '.join(line_parts) + '\n'
                        modified_lines.append(modified_line)
                        continue
            # 不需要修改的行直接保留
            modified_lines.append(line)
        
        # 第三步:把修改后的内容写回文件
        with open(file_path, 'w', encoding='utf-8') as f:
            f.writelines(modified_lines)
        print(f"{filename} 处理完成")

print("所有文件处理完毕!")

注意事项

  • 如果你的文件用制表符分隔,把代码里的split()改成split('\t'),' '.join改成'\t'.join
  • 确保安装了Python 3.x版本,直接运行脚本即可
  • 路径填写要准确,Windows系统注意用双反斜杠转义,Linux/Mac用正斜杠

内容的提问来源于stack exchange,提问作者bioiinffun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 08:38:15