You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Bash的sed、awk提取HTML article标签内容生成文件

问题解答

结论

该需求完全可以在Bash中使用sed、awk工具实现,无需额外安装重型的HTML解析工具即可完成提取、生成文件的需求。

实现说明

你给出的输入HTML中img标签为懒加载格式暂无资源地址,以下脚本默认预留了懒加载属性的提取位置,你可以根据实际的HTML属性调整对应提取规则,完整实现脚本如下:

#!/bin/bash
# 自定义配置项
SITE_DOMAIN="http://example.no"
IMG_DOMAIN="http://imgs.example.no"
GENERATE_TIME=$(date +"%Y-%m-%d %H:%M")

# 用awk按article块拆分提取内容
awk '
BEGIN {
    RS = "</article>"
    file_index = 1
}
/<article/ {
    # 提取新闻链接
    if (match($0, /href="(\/nyheter\/[0-9]+\/)"/, arr)) {
        news_link = "'"$SITE_DOMAIN"'" arr[1]
    }
    # 提取新闻标题,去除多余空白字符
    if (match($0, /<h2[^>]*>([^<]+)<\/h2>/, arr)) {
        news_title = arr[1]
        gsub(/[\n\t ]+/, " ", news_title)
    }
    # 提取图片地址,这里默认提取懒加载常用的data-src属性,有实际src替换属性名即可
    if (match($0, /data-src="([^"]+)"/, arr)) {
        img_link = "'"$IMG_DOMAIN"'" arr[1]
    }
    # 写入对应文件
    output_file = "news" file_index ".txt"
    print news_link > output_file
    print news_title > output_file
    print img_link > output_file
    print "'"$GENERATE_TIME"'" > output_file
    close(output_file)
    file_index++
}
' 你的输入文件名.html

# 批量处理HTML实体乱码
for f in news*.txt; do
    # 可安装recode工具直接转码:recode html..utf8 "$f"
    # 也可手动替换已知乱码:
    sed -i '
        s/&Atilde;&yen;/å/g
        s/&acirc;&#128;&#147;/–/g
    ' "$f"
done

补充说明

  • 脚本执行后会自动按顺序生成news1.txt、news2.txt、news3.txt三个文件,内容格式和你给出的示例完全匹配
  • 如果你输入的HTML中img标签已经有正常src属性,只需要把脚本中data-src替换为src即可正常提取图片地址
  • 生成时间直接调用了系统当前时间,格式符合要求

内容的提问来源于stack exchange,提问作者mohsinali1317

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 04:54:05