You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python文本提取:找到目标起始段后如何定位终止点?

Solution for Extracting Target Segments from Large Text Files

Great job getting to the point of finding your starting lines—that's half the battle! Since the separator line is used for both relevant and irrelevant sections, we just need to locate the first separator that immediately follows each target starting line. Here are two approaches tailored for your 500k-line file:

Basic Approach (Straightforward for Beginners)

First, define your separator as a constant to avoid typos:

SEPARATOR = '-------------------'

Then, iterate through each starting index you found, and scan forward until you hit the separator:

# Store pairs of (start_index, end_index) for each target segment
target_segments = []

for start_idx in start:
    # Start checking from the line right after the target header
    for end_idx in range(start_idx + 1, len(lines)):
        # Use strip() to account for any trailing newlines or spaces in the separator line
        if lines[end_idx].strip() == SEPARATOR:
            target_segments.append( (start_idx, end_idx) )
            break
    else:
        # Handle cases where no separator is found after a start line
        print(f"Warning: No separator found after starting line {start_idx}")

Once you have the segment boundaries, extract the relevant content (the lines between the header and separator):

extracted_info = []
for s_idx, e_idx in target_segments:
    # Skip the header line (s_idx) and stop before the separator (e_idx)
    relevant_lines = lines[s_idx + 1 : e_idx]
    extracted_info.extend(relevant_lines)  # Merge all lines into one list
    # Or keep segments separate: extracted_info.append(relevant_lines)

Optimized Approach (Faster for Large Files)

For 500k lines, scanning forward from each start line can add up. Instead, pre-collect all separator indices first, then use binary search to find the matching end index quickly:

import bisect

# First, get all indices where the separator appears
all_separator_indices = [
    i for i, line in enumerate(lines) 
    if line.strip() == SEPARATOR
]

target_segments = []
for start_idx in start:
    # Find the first separator index that's greater than the start index
    # bisect_left gives us the insertion point, which is our target separator
    separator_pos = bisect.bisect_left(all_separator_indices, start_idx)
    
    if separator_pos < len(all_separator_indices):
        end_idx = all_separator_indices[separator_pos]
        target_segments.append( (start_idx, end_idx) )
    else:
        print(f"Warning: No separator found after starting line {start_idx}")

Binary search cuts the time complexity from O(n) per start line to O(log m) (where m is the number of separators), which makes a huge difference for large datasets.

Key Notes to Avoid Issues

  • Double-check that SEPARATOR matches exactly what's in your file (including any leading/trailing spaces—use strip() only if those aren't part of the actual separator).
  • The else clause in the loop catches cases where a target header doesn't have a matching separator, so you can debug missing data.
  • Since your start list is generated from the file in order, the separator indices will also be in order, so binary search works perfectly.

内容的提问来源于stack exchange,提问作者raptorisa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:38:21