Ubuntu 20.04下用sed删除Price值低于950的文件行求助
删除Price值低于950的行(Ubuntu 20.04环境)
问题背景
待处理的输入文件内容如下:
"addressString":"12366 NY","eId":"64174f8e42b7fdfb837f68b","hasImage":false,"Price":5800,"Name":Bernard Bernoulli,"headline":"nice Fiat 500, red, slight damage to left mirror" "addressString":"451 Citadel","eId":"sd3448e42b7368b","year":1976,"hasImage":true,"Price":12220,"Name":Edward Diego,"headline":"Mercedes SLX, no issues" "addressString":"1321 Bejing","eId":"3102ffdb837fssdff3","Price":350,"Name":Jet Li,"headline":"Dodge Viper, no engine, no tires, no windshield; only cash"
需求:删除所有"Price"字段值低于950的行,且每行的字段数量、位置不固定。
原尝试的sed命令未生效:
sed '/"Price":([0-9]|[1-9][0-9]|[1-8][0-9]{2}|9[0-4][0-9]|950),/d' <inputfile >outputfile
原命令失效原因
- sed默认使用基础正则表达式(BRE),
()和{}需要转义,否则会被当作普通字符解析,无法实现分组和重复匹配 - 正则包含了
950,但需求是删除低于950的行,不应包含该值 - 正则末尾的
,限制了Price字段必须在行中间,但如果Price是行内最后一个字段(或后面无逗号),会匹配失败
可行解决方案
方案1:修正sed命令
使用扩展正则表达式(-E参数),调整匹配逻辑覆盖所有情况:
sed -E '/"Price":([0-9]{1,2}|[1-8][0-9]{2}|9[0-4][0-9])(,|$)/d' inputfile > outputfile
参数说明:
-E:启用扩展正则表达式,无需转义()和{}[0-9]{1,2}:匹配1-99的价格值[1-8][0-9]{2}:匹配100-899的价格值9[0-4][0-9]:匹配900-949的价格值(,|$):匹配Price值后的逗号或行尾,兼容字段在任意位置的情况
方案2:用awk更直观处理
awk擅长处理字段值判断,写法更清晰,无需复杂正则:
awk -F'"' '{ for(i=1; i<=NF; i++){ if($i == "Price"){ price = $(i+2) if(price + 0 >= 950){ print break } } } }' inputfile > outputfile
逻辑说明:
-F'"':以双引号作为字段分隔符- 遍历所有字段,找到
"Price"后,提取其后的数值(分隔后$(i+2)对应价格值) price + 0将字符串转为数字,判断是否≥950,满足则打印整行
也可以用更简洁的awk写法:
awk '/"Price":[0-9]+/{ price = substr($0, index($0, "\"Price\":") + 8) price = substr(price, 1, index(price, /[,"]/) - 1) if(price + 0 >= 950) print }' inputfile > outputfile
内容的提问来源于stack exchange,提问作者fonzman
相关产品推荐
相关产品推荐

