You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XML中<file:write>标签path/Path字段提取至CSV的优化方法问询

问题

我需要在XML文件中查找包含<file:write>的标签,这类标签里有path或Path属性。想把文件名和该属性的值提取到CSV文件里,但path/Path字段在标签里的位置不固定,用cut命令搞不定,求更好的实现方法。

当前执行的命令及输出:

find /opt/mortagage/application.xml -type f -exec egrep -ri "&lt;file:write" /dev/null {} + |uniq| sed '/&lt;!--.*--&gt;/d' | sed '/&lt;!--/,/--&gt;/d'

输出结果:

/opt/mortagage/application.xml:              &lt;file:write doc:id=&quot;16630&quot; path=&quot;${file.location}&quot; doc:name=&quot;Save file to directory&quot;&gt;
/opt/mortagage/application.xml:                      &lt;file:write doc:name=&quot;Write to complete folder&quot; doc:id=&quot;18890&quot; path='#[&quot;${file.completeLocation}&quot; ++ vars.zipFileName]' config-ref=&quot;File_Config_completed&quot;&gt;
/opt/mortagage/application.xml:                      &lt;file:write doc:name=&quot;Write to complete folder&quot; doc:id=&quot;19990&quot; Path='#[&quot;${file.completeLocation}&quot; ++ vars.zipFileName]' config-ref=&quot;File_Config_completed&quot;&gt;
解决方案

方法1:awk正则匹配(最灵活)

试试用awk直接提取,它能精准匹配不固定位置的path/Path属性,还兼容单双引号包裹的值:

find /opt/mortagage/application.xml -type f -exec egrep -ri "<file:write" /dev/null {} + | uniq | sed '/<!--.*-->/d' | sed '/<!--/,/-->/d' | awk '
match($0, /^([^:]+):.*(path|Path)=([\'"])(.*?)\3/i, arr) {
    print arr[1] "," arr[4]
}' > output.csv
  • 正则里的i表示忽略大小写,能同时匹配path和Path
  • ([\'"])(.*?)\3用来捕获属性值,不管是单引号还是双引号包裹的内容都能准确提取
  • 最后直接输出文件名,属性值的CSV格式,重定向到output.csv即可

方法2:sed替换提取(轻量快捷)

如果更习惯sed,用扩展正则也能搞定:

find /opt/mortagage/application.xml -type f -exec egrep -ri "<file:write" /dev/null {} + | uniq | sed '/<!--.*-->/d' | sed '/<!--/,/-->/d' | sed -E '
s/^([^:]+):.*(path|Path)=([\'"])(.*?)\3.*/\1,\4/i
' > output.csv
  • -E打开扩展正则,i忽略大小写匹配属性名
  • 把整行内容替换成文件名,属性值,多余的标签内容直接丢弃

方法3:XML专用工具(最可靠)

如果XML结构复杂,正则容易踩坑,推荐用xmllint(需要先装libxml2-utils包)来正经解析XML:

# 注意:如果file是命名空间,需要先声明,比如加上 --xpath 'declare namespace file="http://你的命名空间URI"; //file:write/@path | //file:write/@Path'
xmllint --xpath '//file:write/@path | //file:write/@Path' /opt/mortagage/application.xml | awk '
BEGIN {print "Filename,PathValue"}
{
    filename="/opt/mortagage/application.xml"
    match($0, /name="(path|Path)" value="(.*)"/i, arr)
    print filename "," arr[2]
}' > output.csv
  • 用XPath直接定位目标属性,完全符合XML语法规则,不会因为标签属性顺序变化出错
  • 如果XML里file是命名空间,记得在XPath里加上命名空间声明,不然会找不到标签

内容的提问来源于stack exchange,提问作者Praveen Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 06:24:56