You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python程序开发需求:按关键词列表提取文本文件分段内容

Python Program to Extract Text Sections by Keywords

Got it, let's build this program to pull out sections from your text file using your keyword list. I'll help refine your initial approach to make it efficient, handle edge cases (like missing keywords or the last section), and output exactly what you need.

Key Fixes & Improvements Over Your Initial Idea

  • Stop re-reading the file over and over: Your current code opens the file once per keyword, which is slow for large files. We'll read it once (or twice, way better than N times) to collect all necessary info.
  • Handle missing keywords gracefully: Any keyword not found in the file will be skipped automatically, no errors.
  • Properly end the last section: The final section should go all the way to the end of the file, not just stop before a non-existent next keyword.

Step-by-Step Implementation

1. First, Gather All Keyword Positions in One Pass

Let's read the file once and note where each keyword first appears (line number plus the keyword itself). We'll store these in a sorted list so we know the order of sections.

def get_keyword_positions(file_path, keywords):
    positions = []
    # Make a copy of keywords so we don't modify the original list
    remaining_keywords = keywords.copy()
    
    with open(file_path, 'r', encoding='utf8') as f:
        for line_num, line in enumerate(f, 1):
            # Check if any remaining keyword is in this line
            for keyword in remaining_keywords:
                if keyword in line:
                    positions.append((line_num, keyword))
                    remaining_keywords.remove(keyword)
                    break
            # Exit early if we've found all keywords
            if not remaining_keywords:
                break
    
    # Sort positions by line number to ensure sections are in order
    positions.sort()
    return positions

2. Calculate End Lines for Each Section

For each keyword's start line, the end line is the line before the next keyword starts. For the last keyword, the end is the very last line of the file.

def get_end_lines(positions, total_lines):
    end_lines = []
    for i in range(len(positions)):
        if i < len(positions) - 1:
            # End at the line before the next keyword starts
            end_lines.append(positions[i+1][0] - 1)
        else:
            # Last section goes to the end of the file
            end_lines.append(total_lines)
    return end_lines

3. Extract and Print Sections

Now we'll pull out each section and print it in your desired format: section title (keyword) plus line numbers and content.

def extract_sections(file_path, keywords):
    # Read all lines first to get total lines and reuse the content
    with open(file_path, 'r', encoding='utf8') as f:
        all_lines = f.readlines()
    total_lines = len(all_lines)
    
    # Get where each keyword is located
    keyword_positions = get_keyword_positions(file_path, keywords)
    
    if not keyword_positions:
        print("No keywords found in the file.")
        return
    
    # Figure out where each section ends
    end_lines = get_end_lines(keyword_positions, total_lines)
    
    # Print each section with proper formatting
    for idx, (start_line, keyword) in enumerate(keyword_positions):
        end_line = end_lines[idx]
        print(f"--- Section: *{keyword}* ---")
        # Loop through lines from start to end (note: all_lines is 0-indexed)
        for line_idx in range(start_line - 1, end_line):
            line_content = all_lines[line_idx].rstrip('\n')
            print(f"Line {line_idx + 1}: {line_content}")
        print("\n")  # Add some space between sections for readability

4. How to Use It

Just call the function with your keyword list and file path:

if __name__ == "__main__":
    list_of_words = ['word1', 'word2', 'word3', 'word4']
    file_txt = "your_text_file.txt"  # Replace with your actual file path
    extract_sections(file_txt, list_of_words)

Quick Customization Tips

  • Case-insensitive matching: If you want to find keywords regardless of case, change the check in get_keyword_positions to if keyword.lower() in line.lower().
  • Multiple occurrences of the same keyword: If you need to capture every instance of a keyword (not just the first), remove the remaining_keywords.remove(keyword) line—but note this will create sections for each occurrence.
  • Exact line matches: If you want to match lines that are exactly the keyword (not just contain it), use if line.strip() == keyword.

内容的提问来源于stack exchange,提问作者RTM99

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 16:42:55