You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas如何仅删除文本中超过2位的数字(保留1-2位数字)

问题描述

使用pandas开展文本数据处理时,需要删除文本中所有位数超过2位的连续数字,同时保留1位、2位长度的数字。
初始测试数据集如下:

customerId                            text
0           1  Hello you should call 46232348
1           2                      What is 42
2           3       Is this a number or 23213
3           4               1 person is there
4           5                    It is 4x4 cm

初始实现代码使用\d+匹配所有长度的连续数字做替换,代码如下:

import pandas as pd
d = {
    "customerId": [1, 2, 3, 4, 5],
    "text": ["Hello you should call 46232348",
             "What is 42",
             "Is this a number or 23213",
             '1 person is there',
             'It is 4x4 cm'],
}
df = pd.DataFrame(data=d)
print(df)
df['text_without_number'] = df['text'].str.replace('\d+', '')

print(df)

代码运行后所有数字都被删除,结果不符合预期:

customerId                            text     text_without_number
0           1  Hello you should call 46232348  Hello you should call 
1           2                      What is 42                What is 
2           3       Is this a number or 23213    Is this a number or 
3           4               1 person is there         person is there
4           5                    It is 4x4 cm              It is x cm

预期的正确处理结果需要保留2位及以下长度的数字:

customerId                            text     text_without_number
0           1  Hello you should call 46232348  Hello you should call 
1           2                      What is 42             What is 42  
2           3       Is this a number or 23213    Is this a number or 
3           4               1 person is there      1 person is there
4           5                    It is 4x4 cm           It is 4x4 cm
实现方案

原代码的问题在于正则规则\d+会匹配所有长度≥1的连续数字,因此会无差别删除所有数字。要实现仅删除3位及以上数字的需求,只需要将正则匹配规则修改为\d{3,},匹配所有连续3个及以上的数字串替换为空即可。
修正后的代码如下:

import pandas as pd
d = {
    "customerId": [1, 2, 3, 4, 5],
    "text": ["Hello you should call 46232348",
             "What is 42",
             "Is this a number or 23213",
             '1 person is there',
             'It is 4x4 cm'],
}
df = pd.DataFrame(data=d)
# 仅替换3位及以上的连续数字,显式声明regex=True避免pandas版本差异导致规则失效
df['text_without_number'] = df['text'].str.replace(r'\d{3,}', '', regex=True)
print(df)

运行后即可得到符合预期的结果:

customerId                            text     text_without_number
0           1  Hello you should call 46232348  Hello you should call 
1           2                      What is 42               What is 42
2           3       Is this a number or 23213    Is this a number or 
3           4               1 person is there       1 person is there
4           5                    It is 4x4 cm            It is 4x4 cm

规则说明

  • \d{3,}是正则表达式的量词写法,代表匹配连续出现至少3次的数字字符,刚好对应「位数超过2位的数字」的匹配需求
  • 正则字符串前加r声明为Python原生字符串,避免转义字符干扰正则规则解析
  • 显式传入regex=True参数,兼容不同pandas版本的replace默认行为,保证代码跨版本稳定运行

内容的提问来源于stack exchange,提问作者Pyyyyt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 22:54:29