如何编写支持跨行插入与删除的正则表达式匹配指定模式?
Hey there! Let's work through your problem step by step. You’ve got three strings: AZZKLMNANAKK, AZZLNAKK, AZLPMNNAK, and you want to build a regex based on AZLN that lets you specify insertion and deletion counts, then find common patterns across all lines. Here’s how to do it:
First, clarify terms
Let’s define what we mean by insertions and deletions relative to your base pattern AZLN:
- Deletions: Removing one or more characters from
AZLN(e.g.,AZLis deleting the finalN,ALNis deleting theZ). - Insertions: Adding any number of extra characters between (or around) the characters of
AZLN(e.g.,AZZLNinserts an extraZbetweenAandZ,AZLPMNinsertsPMbetweenLandN).
Step 1: Build a regex for configurable deletions
To allow up to d deletions, we need to match all valid subsequences of AZLN that retain at least 4 - d characters (keeping their original order).
For example, if you allow 1 deletion max (so we can retain 3 or 4 characters from AZLN), the valid subsequences are:AZLN, AZL, AZN, ALN, ZLN
We can turn these into regex fragments by adding optional gaps for insertions next.
Step 2: Add configurable insertion support
If you allow up to m insertions between each base character, we use .{0,m} (matches 0 to m of any character) between each part of the subsequence.
Combining with the 1-deletion example and 2 insertions max per gap, the regex becomes:
A.{0,2}Z.{0,2}L.{0,2}N|A.{0,2}Z.{0,2}L|A.{0,2}Z.{0,2}N|A.{0,2}L.{0,2}N|Z.{0,2}L.{0,2}N
Step 3: Find common patterns across all lines
To get patterns that exist in every string, follow these steps:
- Use the generated regex to extract all matching substrings from each of your lines.
- Find the intersection of these substring sets — that’s your common patterns.
For your example strings with d=0 (no deletions allowed) and m=3 (max 3 insertions per gap):
- Regex:
A.{0,3}Z.{0,3}L.{0,3}N - Matches from each line:
AZZKLN(first line),AZZLN(second),AZLPMN(third) - The smallest common pattern here is
AZLNitself, since all lines contain this exact sequence in order.
Dynamic regex generation (for flexibility)
Instead of writing regex manually for every combination of deletions/insertions, use a script to generate it automatically. Here’s a Python example:
from itertools import combinations base_pattern = "AZLN" max_deletions = 1 # Adjust to your needs max_insertions_per_gap = 2 # Adjust to your needs # Generate all valid subsequences (length from 4 - max_deletions to 4) valid_subsequences = [] for length in range(len(base_pattern) - max_deletions, len(base_pattern) + 1): for combo in combinations(base_pattern, length): valid_subsequences.append(''.join(combo)) # Build regex fragments with insertion support regex_fragments = [] for seq in valid_subsequences: # Add insertion gaps between each character fragment = ''.join([f'{char}.{{0,{max_insertions_per_gap}}}' for char in seq[:-1]]) + seq[-1] regex_fragments.append(fragment) # Combine into final regex final_regex = '|'.join(regex_fragments) print(final_regex)
Run this script, tweak the max_deletions and max_insertions_per_gap values, and you’ll get a regex tailored to your needs.
Quick notes
- If your strings include newlines, add the single-line mode prefix
(?s)to your regex (e.g.,(?s)A.{0,2}Z...), so.matches line breaks too. - If you need to limit total insertions across the entire pattern (not per gap), use a positive lookahead like
(?=(?:[^AZLN]){0,total_insertions})to count non-base characters.
内容的提问来源于stack exchange,提问作者user2806363

