如何读取文本文件:遇特定值暂停存储后继续读取
问题描述
我有一个如下格式的文本文件:
0 -82.871 2.52531 36.64 138 96.05 0 -76.1014 2.52577 35.36 137 83.9 0 -76.1869 5.57562 35.36 137 62.8 0 -18.1623 -11.6886 386.08 411 200.9 0 -4.62234 -4.91846 325.92 364 252.2 0 -2.52609 -1.63149 325.92 364 85.4 0 -2.52609 -1.63149 112.16 197 48.4 0 -18.1623 -4.91846 -54.24 67 69.55 0 -18.1623 -4.91846 386.08 411 64.55 12345678 1 12345678 2 2 25.2279 -72.3226 48.16 147 221.55 2 28.7109 -70.2263 48.16 147 1587.7 2 76.1009 -63.4562 46.88 146 110.35 2 31.9979 -65.5526 48.16 147 1601.8 2 35.4805 -63.4559 48.16 147 310.25 2 31.9979 -58.7826 49.44 148 492.8 2 35.4805 -56.6859 46.88 146 42.6 2 1.63117 -43.1461 73.76 167 54.55 2 4.91818 -38.4723 76.32 169 75.4
我编写了一段程序,通过line = raw_dat.readlines()[7:]跳过全部表头,读取文件直至遇到magic_number时终止循环:
file = 'Runnumber169raw10.txt' magic_number = '12345678' event1 = [] x1 = [] y1 = [] z1 = [] tb1 = [] q1 = [] Xnoselection = [] X = [] distanceradius = 0 with open(file, 'r') as raw_dat: line = raw_dat.readlines()[7:] for lines in line: lines.split() print(lines) if lines.split()[0] == magic_number: break
break语句会终止循环,此时首列值为0的数据已被存储,但我无法在break后继续读取文件内容,请问如何实现遇特定值时暂停存储、之后继续读取并重复该逻辑?
解决方案
不要用break终止整个循环,而是用状态标记控制存储逻辑,同时持续遍历所有行。核心思路:
- 定义布尔变量(如
should_store),初始设为True,表示默认需要存储数据 - 遍历每一行时,先判断首列是否为
magic_number:- 若是,切换
should_store的状态(True变False或反之),跳过当前行的存储 - 若不是,根据
should_store的状态决定是否存入数据
- 若是,切换
另外注意:原代码中lines.split()未赋值给变量,等于无效操作,需将拆分后的数据存入变量(如parts = lines.split())再使用。
修改后的代码示例:
file = 'Runnumber169raw10.txt' magic_number = '12345678' # 用字典分组存储不同段的数据,比单独定义列表更灵活 data_groups = {} current_group = 'group_0' data_groups[current_group] = {'x': [], 'y': [], 'z': [], 'tb': [], 'q': []} should_store = True with open(file, 'r') as raw_dat: # 跳过前7行表头 for _ in range(7): raw_dat.readline() # 逐行读取剩余内容 for line in raw_dat: line = line.strip() if not line: continue # 跳过空行 parts = line.split() if parts[0] == magic_number: # 遇到魔术数字,切换存储状态,同时准备新分组(如果需要) should_store = not should_store if should_store: # 恢复存储时创建新分组,命名规则可根据需求调整 current_group = f'group_{len(data_groups)}' data_groups[current_group] = {'x': [], 'y': [], 'z': [], 'tb': [], 'q': []} continue # 根据状态决定是否存储数据 if should_store: data_groups[current_group]['x'].append(float(parts[1])) data_groups[current_group]['y'].append(float(parts[2])) data_groups[current_group]['z'].append(float(parts[3])) data_groups[current_group]['tb'].append(int(parts[4])) data_groups[current_group]['q'].append(float(parts[5])) # 验证结果(仅打印每组前3个数据示例) for group, values in data_groups.items(): print(f"\n{group} 数据:") print(f"x: {values['x'][:3]}...") print(f"y: {values['y'][:3]}...")
关键改进点
- 用状态标记替代
break,实现"暂停/恢复"存储的循环逻辑 - 采用字典分组存储,避免定义大量独立列表,扩展性更强
- 逐行读取文件,内存占用更低,适合处理大文件
- 修复了原代码中
split()未赋值的无效操作问题
内容的提问来源于stack exchange,提问作者theheretic
相关产品推荐
相关产品推荐

