求BASH的sed/grep脚本提取含非法注释的指定行内容
问题
我有一个用于移动端自动化测试的JavaScript映射文件,页面元素定义示例如下:
shop_myBoutique: { self: 'android=new UiSelector().className("android.widget.TextView").textContains("My Boutique")' // self: },
行尾的// self:属于非法内联注释,会导致运行时socket错误。需要编写Shell脚本,解析包含数千个元素定义的映射文件,提取所有带有该类非法注释的行,输出从开头的self:到结尾的// self:的完整内容(包含首尾的self:)。
期望输出示例:
self: 'android=new UiSelector().className("android.widget.TextView").textContains("My Boutique")' // self:
我尝试了多个sed和grep正则命令,但均失败($testComments变量存储了整个映射文件内容),尝试的命令如下:
# echo "$testComments" | sed -n '/self: /,/self: /p' > myComments.txt # echo "$testComments" | sed -e "s/.*self: '\([^']*\)'>.*self: /\1/p" > myComments.txt # echo "$testComments" | sed -n "/self:/,/self:/p" > myComments.txt # sed -n "/self: /,/self: /p" > myComments.txt # grep -o '(?<=self) (?s).*(?=self)' $testComments > myComments.txt # echo "$testComments" | sed -e 's/self: \(.*\)\/\/ self:/\1/' | grep '.*self' > myComments.txt # echo "$testComments" | grep '.*self' > myComments.txt # echo "$testComments" | sed '([self: ]+).*([self: ]+)' | grep '.*self:' # echo "$testComments" | sed 's/\(^self:).*( self: $)' > myComments.txt # echo "$testComments" | sed '/self: /,/'\\ self: /p' > myComments.txt echo "$testComments" | grep '(?<=(self: )).*(?= self:)' | grep -o '^[^ ]*' > myComments.txt
解决方案
方法1:使用grep命令
直接用grep的精准匹配提取目标内容,适合简单场景:
# 从变量读取内容 echo "$testComments" | grep -o 'self: '\''.*'\'' // self:' > myComments.txt # 直接读取文件(更高效,适合大文件) grep -o 'self: '\''.*'\'' // self:' your-mapping-file.js > myComments.txt
说明:
-o参数指定只输出匹配到的部分,自动忽略行首的缩进和其他无关内容- 正则
self: '\''.*'\'' // self:'精准匹配从self: '开始,到' // self:结束的完整内容,其中\'是对单引号的转义,确保正则能正确识别字符串中的单引号。
方法2:使用sed命令
如果需要后续扩展处理(比如批量删除这些注释),sed的灵活性更强:
# 从变量读取内容 echo "$testComments" | sed -n 's/.*\(self: '\''.*'\'' // self:\).*/\1/p' > myComments.txt # 直接读取文件 sed -n 's/.*\(self: '\''.*'\'' // self:\).*/\1/p' your-mapping-file.js > myComments.txt
说明:
-n表示只输出经过处理匹配的行- 正则
.*\(self: '\''.*'\'' // self:\).*用括号捕获目标内容,\1提取捕获的部分并输出,自动过滤行首缩进和行尾其他冗余字符。
失败原因分析
之前的命令大多因为以下问题失效:
- 部分命令使用了跨行匹配语法(如
(?s)),但grep默认不支持,且本场景不需要跨行匹配 - 正则中的转义处理错误,比如单引号、斜杠的转义未正确处理,导致无法匹配目标内容
- 过度使用复杂的前后断言,反而降低了匹配的稳定性,本场景用简单的精准匹配即可满足需求
内容的提问来源于stack exchange,提问作者Wulf
相关产品推荐
相关产品推荐

