You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在大文本文件中通过正则表达式匹配含缺失词(下划线占位)的目标语句

Hey Oskar, let's fix this step by step since you're new to regex. Here's a complete solution tailored to your needs, with clear explanations so you understand every part:

Step 1: Full Working Code

import re

# Your target pattern list
message_list = ["das", "_", "mir", "_", "_", "_", "vor"]

# Read the book line by line (tracks line numbers)
def read_book(file_path):
    with open(file_path, 'r', encoding='utf-8') as f:
        # Return list of (line_number, line_content) pairs (line numbers start at 1)
        return [(i+1, line.strip()) for i, line in enumerate(f)]

# Convert your message_list into a regex pattern
def build_pattern(message_list):
    pattern_parts = []
    for item in message_list:
        if item == "_":
            # Match any German word (includes umlauts and ß)
            pattern_parts.append(r'[\wäöüßÄÖÜ]+')
        else:
            # Match exact fixed words, with word boundaries to avoid partial matches
            pattern_parts.append(r'\b' + re.escape(item) + r'\b')
    # Join parts with one or more spaces (handles multiple spaces between words)
    return r'\s+'.join(pattern_parts)

# Main logic to find matches
book_lines = read_book("Book.txt")
target_pattern = build_pattern(message_list)
matches = []

for line_num, line_content in book_lines:
    # Check if the line contains the pattern
    match_result = re.search(target_pattern, line_content)
    if match_result:
        # Extract context around the match for readability
        match_text = match_result.group()
        start_idx = max(0, match_result.start() - 20)
        end_idx = min(len(line_content), match_result.end() + 20)
        context = line_content[start_idx:end_idx]
        # Add ellipses if we cut off text
        if start_idx > 0:
            context = "..." + context
        if end_idx < len(line_content):
            context = context + "..."
        matches.append((line_num, context))

# Print results
print(f"Target Pattern: {message_list}")
if matches:
    print("Found matching lines:")
    for line_num, context in matches:
        print(f"Line {line_num}: {context}")
else:
    print("No matches found in the book.")

Step 2: Key Explanations (For Regex Newbies)

Let's break down the most important parts so you know how this works:

Reading the Book

  • The read_book function reads your Book.txt line by line and keeps track of line numbers (starting at 1, which is more natural for reading books). This lets us report exactly where matches are found.
  • We use encoding='utf-8' to handle German special characters (ä, ö, ü, ß) without weird garbled text.

Building the Regex Pattern

This is the core part you struggled with:

  • For fixed words like "das" or "mir":
    • We wrap them in \b (word boundary) to ensure we match the full word, not just a part of a longer word (e.g., we won't accidentally match "mir" in a made-up word like "mirror").
    • re.escape() makes sure if your fixed words ever have regex special characters (like . or *), they're treated as normal text instead of regex commands.
  • For underscores "_":
    • We use [\wäöüßÄÖÜ]+ to match any valid German word:
      • [] = match any character inside the brackets
      • \w = match standard letters/numbers
      • äöüßÄÖÜ = include German-specific characters
      • + = match at least one character (so we get full words, not empty spaces)
  • We join all parts with \s+, which matches one or more spaces (handles cases where there are multiple spaces between words).

Finding and Displaying Matches

  • We loop through every line of the book, using re.search() to check if the line contains our pattern.
  • When we find a match, we extract a small context around it (20 characters before and after) and add ellipses (...) if we cut off text, so you can see where the match fits in the sentence.

Example Output

If your book has a line like Das gefällt mir nicht ganz viel zu weit vor!, the output would look like:

Target Pattern: ["das", "_", "mir", "_", "_", "_", "vor"]
Found matching lines:
Line 127: ...Das gefällt mir nicht ganz viel zu weit vor!...

内容的提问来源于stack exchange,提问作者Oskar_GER

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 13:32:33