pandas如何仅删除文本中超过2位的数字(保留1-2位数字)
问题描述
使用pandas开展文本数据处理时,需要删除文本中所有位数超过2位的连续数字,同时保留1位、2位长度的数字。
初始测试数据集如下:
customerId text 0 1 Hello you should call 46232348 1 2 What is 42 2 3 Is this a number or 23213 3 4 1 person is there 4 5 It is 4x4 cm
初始实现代码使用\d+匹配所有长度的连续数字做替换,代码如下:
import pandas as pd d = { "customerId": [1, 2, 3, 4, 5], "text": ["Hello you should call 46232348", "What is 42", "Is this a number or 23213", '1 person is there', 'It is 4x4 cm'], } df = pd.DataFrame(data=d) print(df) df['text_without_number'] = df['text'].str.replace('\d+', '') print(df)
代码运行后所有数字都被删除,结果不符合预期:
customerId text text_without_number 0 1 Hello you should call 46232348 Hello you should call 1 2 What is 42 What is 2 3 Is this a number or 23213 Is this a number or 3 4 1 person is there person is there 4 5 It is 4x4 cm It is x cm
预期的正确处理结果需要保留2位及以下长度的数字:
customerId text text_without_number 0 1 Hello you should call 46232348 Hello you should call 1 2 What is 42 What is 42 2 3 Is this a number or 23213 Is this a number or 3 4 1 person is there 1 person is there 4 5 It is 4x4 cm It is 4x4 cm
实现方案
原代码的问题在于正则规则\d+会匹配所有长度≥1的连续数字,因此会无差别删除所有数字。要实现仅删除3位及以上数字的需求,只需要将正则匹配规则修改为\d{3,},匹配所有连续3个及以上的数字串替换为空即可。
修正后的代码如下:
import pandas as pd d = { "customerId": [1, 2, 3, 4, 5], "text": ["Hello you should call 46232348", "What is 42", "Is this a number or 23213", '1 person is there', 'It is 4x4 cm'], } df = pd.DataFrame(data=d) # 仅替换3位及以上的连续数字,显式声明regex=True避免pandas版本差异导致规则失效 df['text_without_number'] = df['text'].str.replace(r'\d{3,}', '', regex=True) print(df)
运行后即可得到符合预期的结果:
customerId text text_without_number 0 1 Hello you should call 46232348 Hello you should call 1 2 What is 42 What is 42 2 3 Is this a number or 23213 Is this a number or 3 4 1 person is there 1 person is there 4 5 It is 4x4 cm It is 4x4 cm
规则说明
\d{3,}是正则表达式的量词写法,代表匹配连续出现至少3次的数字字符,刚好对应「位数超过2位的数字」的匹配需求- 正则字符串前加
r声明为Python原生字符串,避免转义字符干扰正则规则解析 - 显式传入
regex=True参数,兼容不同pandas版本的replace默认行为,保证代码跨版本稳定运行
内容的提问来源于stack exchange,提问作者Pyyyyt
相关产品推荐
相关产品推荐

