如何用sed正确处理含多点的HTML文件名 生成指定h1标签插入文件首行
问题背景
现有一批HTML格式的man手册文件,文件名如下:
ati.4.html fbdevhw.4.html isdn.ctrl.4.html modul.efile.4.html ran.dom.4m.html tw.policy.4p.html
需要给每个文件的第一行插入对应格式的h1标题,目标标题格式为:
<h1>ati(4) - some text tp append</h1> <h1>fbdevhw(4) - some text tp append</h1> <h1>isdn.ctrl(4) - some text tp append</h1> <h1>modul.efile(4) - some text tp append</h1> <h1>ran.dom(4m) - some text tp append</h1> <h1>tw.policy(4p) - some text tp append</h1>
目前已编写shell循环配合sed的脚本,接近预期效果,但存在匹配错误,现有代码如下:
for filename in `ls` do rep_text=`echo $filename | sed 's/\.html/\) - some text tp append<\/h1>/' | sed 's/^/<h1>/' sed -i "1 i\${rep_text}" $filename done
该脚本实际生成的标题存在格式错误,输出如下:
<h1>ati.4) - some text tp append</h1> <h1>fbdevhw.4) - some text tp append</h1> <h1>isdn.ctrl.4) - some text tp append</h1> <h1>modul.efile.4) - some text tp append</h1> <h1>ran.dom.4m) - some text tp append</h1> <h1>tw.policy.4p) - some text tp append</h1>
核心问题:文件名可能包含多个点号,原有逻辑无法精准定位紧邻.html后缀的最后一个分段点,无法将该点正确替换为左括号(,导致标题格式不符合要求,同时希望尽量简化实现逻辑。
实现方案
核心解决思路:从文件名后缀往前做正则匹配,而不是从开头数点的位置。所有目标文件名的结构统一为[手册名(可含点)].[章节号].html,用正则捕获组分别提取手册名和章节号,再拼接为目标标题即可,完全不受文件名内点数量的影响。
修正后易读的循环实现(推荐)
避开原脚本for ls的坑(遇含空格文件名会出错),直接用通配符匹配html文件,单次sed即可完成文件名到标题的转换:
# 遍历当前目录所有html文件 for file in *.html; do # 正则替换生成目标标题:捕获最后一个点前的手册名、点和html间的章节号 title=$(echo "$file" | sed 's/\(.*\)\.\([^.]*\)\.html$/<h1>\1(\2) - some text tp append<\/h1>/') # 将标题插入文件第一行 sed -i "1i $title" "$file" done
正则逻辑说明:
\(.*\)\.:贪婪匹配最后一个点之前的所有内容,作为手册名存入第一个捕获组\([^.]*\)\.html$:匹配最后一个点到.html后缀之间不含点的内容(即章节号,如4、4m、4p),存入第二个捕获组- 替换时直接拼接为
<h1>手册名(章节号) - 固定文本</h1>即可,完全适配含多点的文件名。
GNU sed 单命令批量实现
如果需要压缩为单条sed命令批量处理所有文件(依赖GNU sed特性),可以使用如下写法:
sed -i ' 1{ h s/.*/cat <<< "&"/e s/\(.*\)\.\([^.]*\)\.html/<h1>\1(\2) - some text tp append<\/h1>/ G s/\n/\n/ } ' *.html
该写法不需要shell循环,单次sed进程即可完成所有文件的插入操作,执行效率更高,但可读性弱于循环写法。
内容的提问来源于stack exchange,提问作者Sandeep_Patil
相关产品推荐
相关产品推荐

