如何在大文本文件中通过正则表达式匹配含缺失词(下划线占位)的目标语句
Hey Oskar, let's fix this step by step since you're new to regex. Here's a complete solution tailored to your needs, with clear explanations so you understand every part:
Step 1: Full Working Code
import re # Your target pattern list message_list = ["das", "_", "mir", "_", "_", "_", "vor"] # Read the book line by line (tracks line numbers) def read_book(file_path): with open(file_path, 'r', encoding='utf-8') as f: # Return list of (line_number, line_content) pairs (line numbers start at 1) return [(i+1, line.strip()) for i, line in enumerate(f)] # Convert your message_list into a regex pattern def build_pattern(message_list): pattern_parts = [] for item in message_list: if item == "_": # Match any German word (includes umlauts and ß) pattern_parts.append(r'[\wäöüßÄÖÜ]+') else: # Match exact fixed words, with word boundaries to avoid partial matches pattern_parts.append(r'\b' + re.escape(item) + r'\b') # Join parts with one or more spaces (handles multiple spaces between words) return r'\s+'.join(pattern_parts) # Main logic to find matches book_lines = read_book("Book.txt") target_pattern = build_pattern(message_list) matches = [] for line_num, line_content in book_lines: # Check if the line contains the pattern match_result = re.search(target_pattern, line_content) if match_result: # Extract context around the match for readability match_text = match_result.group() start_idx = max(0, match_result.start() - 20) end_idx = min(len(line_content), match_result.end() + 20) context = line_content[start_idx:end_idx] # Add ellipses if we cut off text if start_idx > 0: context = "..." + context if end_idx < len(line_content): context = context + "..." matches.append((line_num, context)) # Print results print(f"Target Pattern: {message_list}") if matches: print("Found matching lines:") for line_num, context in matches: print(f"Line {line_num}: {context}") else: print("No matches found in the book.")
Step 2: Key Explanations (For Regex Newbies)
Let's break down the most important parts so you know how this works:
Reading the Book
- The
read_bookfunction reads yourBook.txtline by line and keeps track of line numbers (starting at 1, which is more natural for reading books). This lets us report exactly where matches are found. - We use
encoding='utf-8'to handle German special characters (ä, ö, ü, ß) without weird garbled text.
Building the Regex Pattern
This is the core part you struggled with:
- For fixed words like
"das"or"mir":- We wrap them in
\b(word boundary) to ensure we match the full word, not just a part of a longer word (e.g., we won't accidentally match "mir" in a made-up word like "mirror"). re.escape()makes sure if your fixed words ever have regex special characters (like.or*), they're treated as normal text instead of regex commands.
- We wrap them in
- For underscores
"_":- We use
[\wäöüßÄÖÜ]+to match any valid German word:[]= match any character inside the brackets\w= match standard letters/numbersäöüßÄÖÜ= include German-specific characters+= match at least one character (so we get full words, not empty spaces)
- We use
- We join all parts with
\s+, which matches one or more spaces (handles cases where there are multiple spaces between words).
Finding and Displaying Matches
- We loop through every line of the book, using
re.search()to check if the line contains our pattern. - When we find a match, we extract a small context around it (20 characters before and after) and add ellipses (
...) if we cut off text, so you can see where the match fits in the sentence.
Example Output
If your book has a line like Das gefällt mir nicht ganz viel zu weit vor!, the output would look like:
Target Pattern: ["das", "_", "mir", "_", "_", "_", "vor"] Found matching lines: Line 127: ...Das gefällt mir nicht ganz viel zu weit vor!...
内容的提问来源于stack exchange,提问作者Oskar_GER
相关产品推荐
相关产品推荐

