Python程序开发需求:按关键词列表提取文本文件分段内容
Got it, let's build this program to pull out sections from your text file using your keyword list. I'll help refine your initial approach to make it efficient, handle edge cases (like missing keywords or the last section), and output exactly what you need.
Key Fixes & Improvements Over Your Initial Idea
- Stop re-reading the file over and over: Your current code opens the file once per keyword, which is slow for large files. We'll read it once (or twice, way better than N times) to collect all necessary info.
- Handle missing keywords gracefully: Any keyword not found in the file will be skipped automatically, no errors.
- Properly end the last section: The final section should go all the way to the end of the file, not just stop before a non-existent next keyword.
Step-by-Step Implementation
1. First, Gather All Keyword Positions in One Pass
Let's read the file once and note where each keyword first appears (line number plus the keyword itself). We'll store these in a sorted list so we know the order of sections.
def get_keyword_positions(file_path, keywords): positions = [] # Make a copy of keywords so we don't modify the original list remaining_keywords = keywords.copy() with open(file_path, 'r', encoding='utf8') as f: for line_num, line in enumerate(f, 1): # Check if any remaining keyword is in this line for keyword in remaining_keywords: if keyword in line: positions.append((line_num, keyword)) remaining_keywords.remove(keyword) break # Exit early if we've found all keywords if not remaining_keywords: break # Sort positions by line number to ensure sections are in order positions.sort() return positions
2. Calculate End Lines for Each Section
For each keyword's start line, the end line is the line before the next keyword starts. For the last keyword, the end is the very last line of the file.
def get_end_lines(positions, total_lines): end_lines = [] for i in range(len(positions)): if i < len(positions) - 1: # End at the line before the next keyword starts end_lines.append(positions[i+1][0] - 1) else: # Last section goes to the end of the file end_lines.append(total_lines) return end_lines
3. Extract and Print Sections
Now we'll pull out each section and print it in your desired format: section title (keyword) plus line numbers and content.
def extract_sections(file_path, keywords): # Read all lines first to get total lines and reuse the content with open(file_path, 'r', encoding='utf8') as f: all_lines = f.readlines() total_lines = len(all_lines) # Get where each keyword is located keyword_positions = get_keyword_positions(file_path, keywords) if not keyword_positions: print("No keywords found in the file.") return # Figure out where each section ends end_lines = get_end_lines(keyword_positions, total_lines) # Print each section with proper formatting for idx, (start_line, keyword) in enumerate(keyword_positions): end_line = end_lines[idx] print(f"--- Section: *{keyword}* ---") # Loop through lines from start to end (note: all_lines is 0-indexed) for line_idx in range(start_line - 1, end_line): line_content = all_lines[line_idx].rstrip('\n') print(f"Line {line_idx + 1}: {line_content}") print("\n") # Add some space between sections for readability
4. How to Use It
Just call the function with your keyword list and file path:
if __name__ == "__main__": list_of_words = ['word1', 'word2', 'word3', 'word4'] file_txt = "your_text_file.txt" # Replace with your actual file path extract_sections(file_txt, list_of_words)
Quick Customization Tips
- Case-insensitive matching: If you want to find keywords regardless of case, change the check in
get_keyword_positionstoif keyword.lower() in line.lower(). - Multiple occurrences of the same keyword: If you need to capture every instance of a keyword (not just the first), remove the
remaining_keywords.remove(keyword)line—but note this will create sections for each occurrence. - Exact line matches: If you want to match lines that are exactly the keyword (not just contain it), use
if line.strip() == keyword.
内容的提问来源于stack exchange,提问作者RTM99

