如何使用Bash的sed、awk提取HTML article标签内容生成文件
问题解答
结论
该需求完全可以在Bash中使用sed、awk工具实现,无需额外安装重型的HTML解析工具即可完成提取、生成文件的需求。
实现说明
你给出的输入HTML中img标签为懒加载格式暂无资源地址,以下脚本默认预留了懒加载属性的提取位置,你可以根据实际的HTML属性调整对应提取规则,完整实现脚本如下:
#!/bin/bash # 自定义配置项 SITE_DOMAIN="http://example.no" IMG_DOMAIN="http://imgs.example.no" GENERATE_TIME=$(date +"%Y-%m-%d %H:%M") # 用awk按article块拆分提取内容 awk ' BEGIN { RS = "</article>" file_index = 1 } /<article/ { # 提取新闻链接 if (match($0, /href="(\/nyheter\/[0-9]+\/)"/, arr)) { news_link = "'"$SITE_DOMAIN"'" arr[1] } # 提取新闻标题,去除多余空白字符 if (match($0, /<h2[^>]*>([^<]+)<\/h2>/, arr)) { news_title = arr[1] gsub(/[\n\t ]+/, " ", news_title) } # 提取图片地址,这里默认提取懒加载常用的data-src属性,有实际src替换属性名即可 if (match($0, /data-src="([^"]+)"/, arr)) { img_link = "'"$IMG_DOMAIN"'" arr[1] } # 写入对应文件 output_file = "news" file_index ".txt" print news_link > output_file print news_title > output_file print img_link > output_file print "'"$GENERATE_TIME"'" > output_file close(output_file) file_index++ } ' 你的输入文件名.html # 批量处理HTML实体乱码 for f in news*.txt; do # 可安装recode工具直接转码:recode html..utf8 "$f" # 也可手动替换已知乱码: sed -i ' s/Ã¥/å/g s/–/–/g ' "$f" done
补充说明
- 脚本执行后会自动按顺序生成
news1.txt、news2.txt、news3.txt三个文件,内容格式和你给出的示例完全匹配 - 如果你输入的HTML中img标签已经有正常
src属性,只需要把脚本中data-src替换为src即可正常提取图片地址 - 生成时间直接调用了系统当前时间,格式符合要求
内容的提问来源于stack exchange,提问作者mohsinali1317
相关产品推荐
相关产品推荐

