Python如何按指定格式分区读取文本文件?
当然可以啦!Python完全能搞定这种按指定案例ID和对应行数分区读取文本文件的需求,我给你两种实用的思路,你可以根据自己的文件情况选:
方法1:逐行扫描定位目标内容
这种方法兼容性拉满,不管文件大小如何、换行符有没有差异都能用,逻辑也特别直观:
- 打开文件后逐行遍历,同时记录当前是否处于目标案例的内容读取状态
- 当碰到
1 100 ---stuff----这类标识行时,判断是不是你要找的案例ID - 如果是目标ID,就开始读取接下来的指定行数;如果不是,直接跳过对应行数
代码示例
def read_case_content(file_path, target_case_ids): target_cases = {} current_case_id = None lines_remaining = 0 content_buffer = [] with open(file_path, 'r', encoding='utf-8') as f: for line in f: stripped_line = line.strip() # 判断是否为案例标识行(格式:案例ID 行数 ---内容标记---) if stripped_line and stripped_line[0].isdigit(): parts = stripped_line.split(maxsplit=2) if len(parts) >= 2 and parts[1].isdigit(): # 先保存上一个已读完的目标案例内容 if current_case_id in target_case_ids and lines_remaining == 0: target_cases[current_case_id] = '\n'.join(content_buffer) content_buffer = [] # 更新当前案例信息 current_case_id = parts[0] lines_remaining = int(parts[1]) # 如果当前案例不是目标,直接跳过后续内容 if current_case_id not in target_case_ids: lines_remaining = 0 else: # 若当前是目标案例且还有内容要读 if lines_remaining > 0: content_buffer.append(stripped_line) lines_remaining -= 1 # 处理文件最后一个案例的内容 if current_case_id in target_case_ids and lines_remaining == 0: target_cases[current_case_id] = '\n'.join(content_buffer) return target_cases # 使用示例 result = read_case_content('your_target_file.txt', ['1', '2']) print("案例1的内容:") print(result.get('1', '未找到对应案例')) print("\n案例2的内容:") print(result.get('2', '未找到对应案例'))
方法2:计算偏移量直接跳转(适合固定结构的大文件)
如果你的文件结构非常规整,每行的字节数稳定、编码统一,那可以用这种效率更高的方法:
- 先扫一遍文件,记录每个案例标识行之后的内容起始偏移量和对应行数
- 之后读取目标案例时,直接用
seek()跳转到对应位置,读取指定行数即可
代码示例
def get_case_position_map(file_path): case_positions = {} with open(file_path, 'rb') as f: while True: line = f.readline() if not line: break try: decoded_line = line.decode('utf-8').strip() parts = decoded_line.split(maxsplit=2) if len(parts) >= 2 and parts[1].isdigit(): case_id = parts[0] content_line_count = int(parts[1]) # 记录内容开始的位置(标识行之后的位置) content_start = f.tell() case_positions[case_id] = (content_start, content_line_count) # 跳过当前案例的内容行,继续扫描下一个案例 for _ in range(content_line_count): f.readline() except UnicodeDecodeError: # 若遇到编码问题,根据实际文件编码调整这里的处理逻辑 pass return case_positions def read_case_by_position(file_path, position_info): content = [] start_pos, line_count = position_info with open(file_path, 'r', encoding='utf-8') as f: f.seek(start_pos) for _ in range(line_count): content.append(f.readline().strip()) return '\n'.join(content) # 使用示例 case_map = get_case_position_map('your_target_file.txt') # 读取案例1 case1_content = read_case_by_position('your_target_file.txt', case_map['1']) print("案例1内容:\n", case1_content) # 读取案例2 case2_content = read_case_by_position('your_target_file.txt', case_map['2']) print("\n案例2内容:\n", case2_content)
小提示
- 方法1适合大多数场景,不用纠结文件的底层结构,容错性好
- 方法2更适合超大文件,读取速度更快,但要求文件结构稳定,编码统一
内容的提问来源于stack exchange,提问作者Han Ben
相关产品推荐
相关产品推荐

