Python中合并数字正则表达式的有效方法咨询
Hey there! I see you've been searching for a way to merge your digit-focused regex patterns and haven't landed on a clear approach yet. Let's break this down step by step—first manually, then with a simple code solution to automate the process, which should align perfectly with your expected output.
Step 1: Group Regexes by Their Common Prefix
The first key step is to group all regex patterns that share the same leading digits (the part before the [...] or final single digit). For your input list:
- Group 1 (prefix
832118):832118[0-3],832118[7-8] - Group 2 (prefix
832119):832119[0-1],832119[4-6],832119[8-9] - Group 3 (prefix
832120):8321206,832120[0-4],832120[8-9]
Step 2: Extract and Normalize Suffixes
For each group, pull out the suffixes (the parts that define the digit ranges or single digits):
- For patterns with
[...], extract the content inside the brackets (e.g.,0-3from832118[0-3]) - For patterns with a single final digit, treat that digit as a single-character range (e.g.,
6from8321206becomes6-6for consistency)
Step 3: Sort and Merge Suffixes
Sort the normalized suffixes by their starting digit, then convert them back to string format—keeping single digits as just the number, and ranges as start-end. Finally, concatenate all these into a single set of brackets attached to the shared prefix.
For example, for the 832120 group:
- Normalized suffixes:
0-4,6-6,8-9 - Sorted and converted back:
0-4,6,8-9 - Merged into one bracket:
[0-468-9] - Final regex:
832120[0-468-9]
Automating the Process with Python
If you want to do this for larger lists, here's a simple Python script that handles all these steps automatically:
# Your input regex list regex_list = [ "832118[0-3]", "832118[7-8]", "832119[0-1]", "832119[4-6]", "832119[8-9]", "8321206", "832120[0-4]", "832120[8-9]" ] # Group regexes by their common prefix prefix_groups = {} for regex in regex_list: if "[" in regex: # Split into prefix and bracket content prefix = regex.split("[")[0] suffix = regex.split("[")[1].rstrip("]") else: # Single final digit: prefix is all except last character prefix = regex[:-1] suffix = regex[-1] # Add to the group if prefix not in prefix_groups: prefix_groups[prefix] = [] prefix_groups[prefix].append(suffix) # Process each group to merge suffixes merged_regexes = [] for prefix, suffixes in prefix_groups.items(): # Normalize all suffixes to (start, end) tuples normalized_suffixes = [] for s in suffixes: if "-" in s: start, end = map(int, s.split("-")) normalized_suffixes.append((start, end)) else: digit = int(s) normalized_suffixes.append((digit, digit)) # Sort suffixes by their starting digit normalized_suffixes.sort() # Convert back to string format merged_suffix_parts = [] for start, end in normalized_suffixes: if start == end: merged_suffix_parts.append(str(start)) else: merged_suffix_parts.append(f"{start}-{end}") # Combine into the final regex merged_regex = f"{prefix}[{''.join(merged_suffix_parts)}]" merged_regexes.append(merged_regex) # Print the result for regex in merged_regexes: print(regex)
Running This Script Will Output:
832118[0-37-8] 832119[0-14-68-9] 832120[0-468-9]
This script works for any similar list of digit regexes—just update the regex_list with your own patterns.
内容的提问来源于stack exchange,提问作者beekeeper

