You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

构建TXT文件解析HTML标签:从URL列表提取<h1>标签内容

批量提取URL页面H1标签内容并导出为TXT文件

我经常遇到这类批量抓取页面核心标题的需求,用Python结合requests和BeautifulSoup就能轻松搞定,给你一套完整的解决方案:

1. 安装依赖工具

首先需要安装两个必备的Python库,用来处理HTTP请求和HTML解析:

pip install requests beautifulsoup4

2. 完整实现代码

下面的代码会遍历你的URL列表,自动抓取每个页面的<h1>标签文本,处理请求异常,最后导出为结构化的TXT文件:

import requests
from bs4 import BeautifulSoup

# 替换成你实际的URL列表
target_urls = [
    "https://example.com/page1",
    "https://example.com/page2",
    # 可以继续添加更多URL
]

# 输出文件路径,可自定义
output_txt_path = "h1_content_results.txt"

# 打开文件并写入内容
with open(output_txt_path, "w", encoding="utf-8") as output_file:
    for index, url in enumerate(target_urls, start=1):
        try:
            # 发送请求,设置超时避免卡顿
            response = requests.get(url, timeout=15)
            # 检查请求是否成功(比如404、500都会触发异常)
            response.raise_for_status()
            
            # 解析HTML页面
            soup = BeautifulSoup(response.text, "html.parser")
            # 查找第一个h1标签(如果页面有多个h1,可改用find_all())
            h1_element = soup.find("h1")
            
            if h1_element:
                # 提取纯文本,strip()去除首尾空格和换行
                h1_text = h1_element.get_text(strip=True)
                # 写入结构化内容,方便后续阅读或解析
                output_file.write(f"序号: {index}\n")
                output_file.write(f"来源URL: {url}\n")
                output_file.write(f"H1内容: {h1_text}\n")
                output_file.write("="*60 + "\n")
            else:
                output_file.write(f"序号: {index}\n")
                output_file.write(f"来源URL: {url}\n")
                output_file.write("⚠️ 未找到H1标签\n")
                output_file.write("="*60 + "\n")
                
        except Exception as error:
            output_file.write(f"序号: {index}\n")
            output_file.write(f"来源URL: {url}\n")
            output_file.write(f"❌ 处理失败: {str(error)}\n")
            output_file.write("="*60 + "\n")

print(f"所有URL处理完成!结果已保存到: {output_txt_path}")

3. 关键细节说明

  • 异常处理:代码会捕获请求超时、页面不存在、网络错误等问题,避免程序中途崩溃,同时在TXT中记录错误信息
  • 文本清洗:get_text(strip=True)会自动去除H1文本中的多余空格、换行和制表符,让内容更整洁
  • 结构化输出:TXT文件中用分隔线、序号、URL标注内容,方便人工阅读或后续脚本解析
  • 多H1场景:如果页面存在多个<h1>标签,可将soup.find("h1")改为soup.find_all("h1"),然后循环提取每个H1的内容

4. 可选:保留H1标签的HTML结构

如果你需要在TXT中保留完整的H1 HTML标签(比如<h1 class="title">Hello World</h1>),只需将提取文本的代码替换为:

h1_html = str(h1_element)
output_file.write(f"H1原始HTML: {h1_html}\n")

内容的提问来源于stack exchange,提问作者user3520363

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:32:28