You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中对字符串型手机号进行标准化处理

Pandas 手机号列标准化处理方案

先看示例数据和已有的处理步骤:

import pandas as pd

data = {'Name': ['John', 'Dom', 'Jack', 'Sam', 'Fred', 'Harvey', 'Toby'],
        'Phone': ['+49(0) 047905356', '(0161) 496 0674', '239.711.3836', '02984 08192', 
        '(0306) 999 0871', '0121x496x0225', '+44047905356']}

df = pd.DataFrame(data)

# 第一步:去除所有非数字字符
df['Phone'] = df['Phone'].replace('\W','', regex=True)

执行这一步后,Phone列的纯数字结果为:

0    490047905356
1     01614960674
2      2397113836
3      0298408192
4     03069990871
5     01214960225
6     44047905356

接下来完成两个标准化需求:

1. 处理带国家码的国际号码

针对原号码以+开头的情况(去特殊字符后以49/44开头),提取本地号码部分:

  • 对于49开头的号码:去掉前三位490,保留后续本地号码
  • 对于44开头的号码:去掉前两位44,保留后续本地号码

用正则替换实现:

# 处理+49开头的国际号码
df['Phone'] = df['Phone'].replace(r'^490', '', regex=True)
# 处理+44开头的国际号码
df['Phone'] = df['Phone'].replace(r'^44', '', regex=True)

2. 为非0开头的号码添加前置0

用正则匹配开头不为0的数字,在其前方补0:

df['Phone'] = df['Phone'].replace(r'^([1-9])', r'0\1', regex=True)

完整整合代码

import pandas as pd

data = {'Name': ['John', 'Dom', 'Jack', 'Sam', 'Fred', 'Harvey', 'Toby'],
        'Phone': ['+49(0) 047905356', '(0161) 496 0674', '239.711.3836', '02984 08192', 
        '(0306) 999 0871', '0121x496x0225', '+44047905356']}

df = pd.DataFrame(data)

# 1. 去除所有非数字字符
df['Phone'] = df['Phone'].replace('\W','', regex=True)

# 2. 处理带国家码的国际号码
df['Phone'] = df['Phone'].replace(r'^490', '', regex=True)
df['Phone'] = df['Phone'].replace(r'^44', '', regex=True)

# 3. 为非0开头的号码补前置0
df['Phone'] = df['Phone'].replace(r'^([1-9])', r'0\1', regex=True)

print(df['Phone'])

执行后的最终Phone列结果:

0    047905356
1    01614960674
2    02397113836
3    0298408192
4    03069990871
5    01214960225
6    047905356
Name: Phone, dtype: object

完全符合需求:带+号的国际号码转为本地0开头格式,原开头非0的号码补全前置0。

内容的提问来源于stack exchange,提问作者Adam Idris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 15:19:32