如何用Python移除字符串中价格数字两侧的双引号?
问题描述
我有若干如下格式的字符串:
"apple cost": "2.78" "orange cost": "12.59" "melone cost": "42.12"
其中数字是可变的,价格范围在0.01到999.99之间,字符串仅水果名称会变化,但格式始终为"fruit name cost": ,这部分无需修改。我希望移除数字两侧的双引号,得到如下结果:
"apple cost": 2.78 "orange cost": 12.59 "melone cost": 42.12
我尝试了如下代码,但双引号并未被移除:
string = '"apple cost": "2.78"' string = string.replace('"apple cost": "([0-9]).([1-9]|[0-9][0-9])"', '"apple cost": ([0-9]).([1-9]|[0-9][0-9])') string = string.replace('"orange cost": "([0-9]).([1-9]|[0-9][0-9])"', '"orange cost": ([0-9]).([1-9]|[0-9][0-9])')
解决方案
你的问题出在str.replace()方法不支持正则表达式匹配,它只会做完全字面替换,所以那些带括号的正则语法根本不会生效,自然无法移除双引号。正确的做法是用Python的re模块,通过正则表达式匹配并替换。
实现代码
import re # 示例输入字符串 input_str = '''"apple cost": "2.78" "orange cost": "12.59" "melone cost": "42.12"''' # 正则匹配规则:捕获键部分和价格部分,替换时去掉价格的引号 pattern = r'(".*? cost":) "(\d{1,3}\.\d{2})"' output_str = re.sub(pattern, r'\1 \2', input_str) print(output_str)
代码说明
r'(".*? cost":) "(\d{1,3}\.\d{2})"':正则规则拆解(".*? cost":):捕获固定格式的键部分(比如"apple cost":),.*?是非贪婪匹配,确保只匹配到cost":为止,适配任意水果名称"(\d{1,3}\.\d{2})":捕获价格部分,\d{1,3}匹配1-3位整数(对应0-999),\.\d{2}匹配小数点后两位(对应01-99),正好覆盖0.01到999.99的价格范围
re.sub(pattern, r'\1 \2', input_str):替换时,用第一个捕获组(键部分)加空格,再加上第二个捕获组(去掉引号的价格),完成目标替换
运行结果
"apple cost": 2.78 "orange cost": 12.59 "melone cost": 42.12
内容的提问来源于stack exchange,提问作者Kunibert
相关产品推荐
相关产品推荐

