You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何按指定格式分区读取文本文件?

当然可以啦!Python完全能搞定这种按指定案例ID和对应行数分区读取文本文件的需求,我给你两种实用的思路,你可以根据自己的文件情况选:

方法1:逐行扫描定位目标内容

这种方法兼容性拉满,不管文件大小如何、换行符有没有差异都能用,逻辑也特别直观:

  • 打开文件后逐行遍历,同时记录当前是否处于目标案例的内容读取状态
  • 当碰到1 100 ---stuff----这类标识行时,判断是不是你要找的案例ID
  • 如果是目标ID,就开始读取接下来的指定行数;如果不是,直接跳过对应行数

代码示例

def read_case_content(file_path, target_case_ids):
    target_cases = {}
    current_case_id = None
    lines_remaining = 0
    content_buffer = []

    with open(file_path, 'r', encoding='utf-8') as f:
        for line in f:
            stripped_line = line.strip()
            # 判断是否为案例标识行(格式:案例ID 行数 ---内容标记---)
            if stripped_line and stripped_line[0].isdigit():
                parts = stripped_line.split(maxsplit=2)
                if len(parts) >= 2 and parts[1].isdigit():
                    # 先保存上一个已读完的目标案例内容
                    if current_case_id in target_case_ids and lines_remaining == 0:
                        target_cases[current_case_id] = '\n'.join(content_buffer)
                        content_buffer = []
                    # 更新当前案例信息
                    current_case_id = parts[0]
                    lines_remaining = int(parts[1])
                    # 如果当前案例不是目标,直接跳过后续内容
                    if current_case_id not in target_case_ids:
                        lines_remaining = 0
            else:
                # 若当前是目标案例且还有内容要读
                if lines_remaining > 0:
                    content_buffer.append(stripped_line)
                    lines_remaining -= 1
        # 处理文件最后一个案例的内容
        if current_case_id in target_case_ids and lines_remaining == 0:
            target_cases[current_case_id] = '\n'.join(content_buffer)
    
    return target_cases

# 使用示例
result = read_case_content('your_target_file.txt', ['1', '2'])
print("案例1的内容:")
print(result.get('1', '未找到对应案例'))
print("\n案例2的内容:")
print(result.get('2', '未找到对应案例'))

方法2:计算偏移量直接跳转(适合固定结构的大文件)

如果你的文件结构非常规整,每行的字节数稳定、编码统一,那可以用这种效率更高的方法:

  • 先扫一遍文件,记录每个案例标识行之后的内容起始偏移量和对应行数
  • 之后读取目标案例时,直接用seek()跳转到对应位置,读取指定行数即可

代码示例

def get_case_position_map(file_path):
    case_positions = {}
    with open(file_path, 'rb') as f:
        while True:
            line = f.readline()
            if not line:
                break
            try:
                decoded_line = line.decode('utf-8').strip()
                parts = decoded_line.split(maxsplit=2)
                if len(parts) >= 2 and parts[1].isdigit():
                    case_id = parts[0]
                    content_line_count = int(parts[1])
                    # 记录内容开始的位置(标识行之后的位置)
                    content_start = f.tell()
                    case_positions[case_id] = (content_start, content_line_count)
                    # 跳过当前案例的内容行,继续扫描下一个案例
                    for _ in range(content_line_count):
                        f.readline()
            except UnicodeDecodeError:
                # 若遇到编码问题,根据实际文件编码调整这里的处理逻辑
                pass
    return case_positions

def read_case_by_position(file_path, position_info):
    content = []
    start_pos, line_count = position_info
    with open(file_path, 'r', encoding='utf-8') as f:
        f.seek(start_pos)
        for _ in range(line_count):
            content.append(f.readline().strip())
    return '\n'.join(content)

# 使用示例
case_map = get_case_position_map('your_target_file.txt')
# 读取案例1
case1_content = read_case_by_position('your_target_file.txt', case_map['1'])
print("案例1内容:\n", case1_content)
# 读取案例2
case2_content = read_case_by_position('your_target_file.txt', case_map['2'])
print("\n案例2内容:\n", case2_content)

小提示

  • 方法1适合大多数场景,不用纠结文件的底层结构,容错性好
  • 方法2更适合超大文件,读取速度更快,但要求文件结构稳定,编码统一

内容的提问来源于stack exchange,提问作者Han Ben

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:08:54