You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将海量手机号前缀转换为正则表达式?求工具或脚本方案

手机号前缀批量转换为正则表达式的解决方案

一、自动转换工具

可以使用专门的前缀压缩类工具处理这类需求,这类工具能自动识别前缀的公共部分,合并重复分支生成最优正则表达式,适合批量处理大量前缀数据:

  • 操作流程:
    1. 提前将手机号前缀按「地区/运营商」分组(避免不同归属的前缀被错误合并)
    2. 将每组前缀导入工具,选择正则输出格式(比如带^开头和[0-9]+$后缀匹配后续号码)
    3. 生成后直接导出分组后的正则与归属对应表

二、编程语言脚本实现

如果需要自定义逻辑(比如特殊前缀处理、特定格式输出),可以用以下脚本实现:

Python 实现

通过Trie树结构合并公共前缀,生成紧凑的正则表达式:

class TrieNode:
    def __init__(self):
        self.children = {}
        self.is_end = False

def build_trie(prefixes):
    root = TrieNode()
    for prefix in prefixes:
        node = root
        for char in prefix:
            if char not in node.children:
                node.children[char] = TrieNode()
            node = node.children[char]
        node.is_end = True
    return root

def trie_to_regex(node):
    if not node.children:
        return ''
    children = []
    for char, child in node.children.items():
        child_regex = trie_to_regex(child)
        # 处理既是前缀终点又有子节点的情况(比如6221本身是前缀,同时有62212、62213)
        if child.is_end and child_regex:
            children.append(f'{char}(?:{child_regex})?')
        else:
            children.append(f'{char}{child_regex}')
    # 单个分支无需分组,多个分支用非捕获组合并
    return children[0] if len(children) == 1 else f'(?:{"|".join(children)})'

def generate_regex(prefixes):
    trie = build_trie(prefixes)
    regex_body = trie_to_regex(trie)
    return f'^{regex_body}[0-9]+$'

# 示例:处理雅加达地区前缀
jakarta_prefixes = ['6221', '62212', '62213']
print(generate_regex(jakarta_prefixes))  # 输出: ^6221(?:2|3)?[0-9]+$
  • 操作步骤:
    1. 读取原始数据,用字典按「地区/运营商」分组存储前缀列表
    2. 遍历每个分组,调用generate_regex生成对应正则
    3. 将结果整理成表格输出

JavaScript 实现

逻辑与Python一致,适合前端或Node.js环境处理:

class TrieNode {
    constructor() {
        this.children = {};
        this.isEnd = false;
    }
}

function buildTrie(prefixes) {
    const root = new TrieNode();
    for (const prefix of prefixes) {
        let node = root;
        for (const char of prefix) {
            if (!node.children[char]) {
                node.children[char] = new TrieNode();
            }
            node = node.children[char];
        }
        node.isEnd = true;
    }
    return root;
}

function trieToRegex(node) {
    if (Object.keys(node.children).length === 0) {
        return '';
    }
    const children = [];
    for (const [char, child] of Object.entries(node.children)) {
        const childRegex = trieToRegex(child);
        if (child.isEnd && childRegex) {
            children.push(`${char}(?:${childRegex})?`);
        } else {
            children.push(`${char}${childRegex}`);
        }
    }
    return children.length === 1 ? children[0] : `(?:${children.join('|')})`;
}

function generateRegex(prefixes) {
    const trie = buildTrie(prefixes);
    const regexBody = trieToRegex(trie);
    return `^${regexBody}[0-9]+$`;
}

// 示例使用
const jakartaPrefixes = ['6221', '62212', '62213'];
console.log(generateRegex(jakartaPrefixes)); // 输出: ^6221(?:2|3)?[0-9]+$

PHP 实现

适合后端PHP环境批量处理:

class TrieNode {
    public $children = [];
    public $isEnd = false;
}

function buildTrie($prefixes) {
    $root = new TrieNode();
    foreach ($prefixes as $prefix) {
        $node = $root;
        $chars = str_split($prefix);
        foreach ($chars as $char) {
            if (!isset($node->children[$char])) {
                $node->children[$char] = new TrieNode();
            }
            $node = $node->children[$char];
        }
        $node->isEnd = true;
    }
    return $root;
}

function trieToRegex($node) {
    if (empty($node->children)) {
        return '';
    }
    $children = [];
    foreach ($node->children as $char => $child) {
        $childRegex = trieToRegex($child);
        if ($child->isEnd && $childRegex) {
            $children[] = "$char(?:$childRegex)?";
        } else {
            $children[] = "$char$childRegex";
        }
    }
    return count($children) === 1 ? $children[0] : '(?:'.implode('|', $children).')';
}

function generateRegex($prefixes) {
    $trie = buildTrie($prefixes);
    $regexBody = trieToRegex($trie);
    return "^$regexBody[0-9]+$";
}

// 示例使用
$jakartaPrefixes = ['6221', '62212', '62213'];
echo generateRegex($jakartaPrefixes); // 输出: ^6221(?:2|3)?[0-9]+$

注意事项

  • 必须按「地区/运营商」分组后再处理,否则不同归属的前缀会被合并,导致正则对应关系混乱
  • 5万条数据量在脚本处理中完全可行,Trie树的时间复杂度为O(n*m)(n为前缀数量,m为前缀平均长度),内存占用也在合理范围内

内容的提问来源于stack exchange,提问作者Paco Casas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 11:31:02