You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中合并数字正则表达式的有效方法咨询

Merging Digit-Based Regular Expressions with Common Prefixes

Hey there! I see you've been searching for a way to merge your digit-focused regex patterns and haven't landed on a clear approach yet. Let's break this down step by step—first manually, then with a simple code solution to automate the process, which should align perfectly with your expected output.

Step 1: Group Regexes by Their Common Prefix

The first key step is to group all regex patterns that share the same leading digits (the part before the [...] or final single digit). For your input list:

  • Group 1 (prefix 832118): 832118[0-3], 832118[7-8]
  • Group 2 (prefix 832119): 832119[0-1], 832119[4-6], 832119[8-9]
  • Group 3 (prefix 832120): 8321206, 832120[0-4], 832120[8-9]

Step 2: Extract and Normalize Suffixes

For each group, pull out the suffixes (the parts that define the digit ranges or single digits):

  • For patterns with [...], extract the content inside the brackets (e.g., 0-3 from 832118[0-3])
  • For patterns with a single final digit, treat that digit as a single-character range (e.g., 6 from 8321206 becomes 6-6 for consistency)

Step 3: Sort and Merge Suffixes

Sort the normalized suffixes by their starting digit, then convert them back to string format—keeping single digits as just the number, and ranges as start-end. Finally, concatenate all these into a single set of brackets attached to the shared prefix.

For example, for the 832120 group:

  • Normalized suffixes: 0-4, 6-6, 8-9
  • Sorted and converted back: 0-4, 6, 8-9
  • Merged into one bracket: [0-468-9]
  • Final regex: 832120[0-468-9]

Automating the Process with Python

If you want to do this for larger lists, here's a simple Python script that handles all these steps automatically:

# Your input regex list
regex_list = [
    "832118[0-3]", "832118[7-8]",
    "832119[0-1]", "832119[4-6]", "832119[8-9]",
    "8321206", "832120[0-4]", "832120[8-9]"
]

# Group regexes by their common prefix
prefix_groups = {}
for regex in regex_list:
    if "[" in regex:
        # Split into prefix and bracket content
        prefix = regex.split("[")[0]
        suffix = regex.split("[")[1].rstrip("]")
    else:
        # Single final digit: prefix is all except last character
        prefix = regex[:-1]
        suffix = regex[-1]
    
    # Add to the group
    if prefix not in prefix_groups:
        prefix_groups[prefix] = []
    prefix_groups[prefix].append(suffix)

# Process each group to merge suffixes
merged_regexes = []
for prefix, suffixes in prefix_groups.items():
    # Normalize all suffixes to (start, end) tuples
    normalized_suffixes = []
    for s in suffixes:
        if "-" in s:
            start, end = map(int, s.split("-"))
            normalized_suffixes.append((start, end))
        else:
            digit = int(s)
            normalized_suffixes.append((digit, digit))
    
    # Sort suffixes by their starting digit
    normalized_suffixes.sort()
    
    # Convert back to string format
    merged_suffix_parts = []
    for start, end in normalized_suffixes:
        if start == end:
            merged_suffix_parts.append(str(start))
        else:
            merged_suffix_parts.append(f"{start}-{end}")
    
    # Combine into the final regex
    merged_regex = f"{prefix}[{''.join(merged_suffix_parts)}]"
    merged_regexes.append(merged_regex)

# Print the result
for regex in merged_regexes:
    print(regex)

Running This Script Will Output:

832118[0-37-8]
832119[0-14-68-9]
832120[0-468-9]

This script works for any similar list of digit regexes—just update the regex_list with your own patterns.

内容的提问来源于stack exchange,提问作者beekeeper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:00:13