Python文本提取:找到目标起始段后如何定位终止点?
Great job getting to the point of finding your starting lines—that's half the battle! Since the separator line is used for both relevant and irrelevant sections, we just need to locate the first separator that immediately follows each target starting line. Here are two approaches tailored for your 500k-line file:
Basic Approach (Straightforward for Beginners)
First, define your separator as a constant to avoid typos:
SEPARATOR = '-------------------'
Then, iterate through each starting index you found, and scan forward until you hit the separator:
# Store pairs of (start_index, end_index) for each target segment target_segments = [] for start_idx in start: # Start checking from the line right after the target header for end_idx in range(start_idx + 1, len(lines)): # Use strip() to account for any trailing newlines or spaces in the separator line if lines[end_idx].strip() == SEPARATOR: target_segments.append( (start_idx, end_idx) ) break else: # Handle cases where no separator is found after a start line print(f"Warning: No separator found after starting line {start_idx}")
Once you have the segment boundaries, extract the relevant content (the lines between the header and separator):
extracted_info = [] for s_idx, e_idx in target_segments: # Skip the header line (s_idx) and stop before the separator (e_idx) relevant_lines = lines[s_idx + 1 : e_idx] extracted_info.extend(relevant_lines) # Merge all lines into one list # Or keep segments separate: extracted_info.append(relevant_lines)
Optimized Approach (Faster for Large Files)
For 500k lines, scanning forward from each start line can add up. Instead, pre-collect all separator indices first, then use binary search to find the matching end index quickly:
import bisect # First, get all indices where the separator appears all_separator_indices = [ i for i, line in enumerate(lines) if line.strip() == SEPARATOR ] target_segments = [] for start_idx in start: # Find the first separator index that's greater than the start index # bisect_left gives us the insertion point, which is our target separator separator_pos = bisect.bisect_left(all_separator_indices, start_idx) if separator_pos < len(all_separator_indices): end_idx = all_separator_indices[separator_pos] target_segments.append( (start_idx, end_idx) ) else: print(f"Warning: No separator found after starting line {start_idx}")
Binary search cuts the time complexity from O(n) per start line to O(log m) (where m is the number of separators), which makes a huge difference for large datasets.
Key Notes to Avoid Issues
- Double-check that
SEPARATORmatches exactly what's in your file (including any leading/trailing spaces—usestrip()only if those aren't part of the actual separator). - The
elseclause in the loop catches cases where a target header doesn't have a matching separator, so you can debug missing data. - Since your
startlist is generated from the file in order, the separator indices will also be in order, so binary search works perfectly.
内容的提问来源于stack exchange,提问作者raptorisa

