You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取第一个h1标签后的文本?网页批量爬取技术方案

问题

每日需爬取并清洗100个网站文本,遇到一类存在多个<h1>标签的网站:滚动至下一个<h1>标签时URL会变化(示例地址:https://economictimes.indiatimes.com/news/international/business/volkswagen-sets-5-7-revenue-growth-target-preaches-cost-discipline/articleshow/101168014.cms)。需提取第一个<h1>标签之后的文本(该文本不在<p>标签内),现有代码逻辑存在问题,请求修正。

现有代码:

response=requests.get('https://economictimes.indiatimes.com/news/international/business/volkswagen-sets-5-7-revenue-growth-target-preaches-cost-discipline/articleshow/101168014.cms',headers={"User-Agent" : "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/51.0.2704.103 Safari/537.36"})      
soup = BeautifulSoup(response.content, 'html.parser')
 if len(soup.body.find_all('h1'))>2:    #to check if there is more than one tag      
    if i.endswith(".cms"):              #to check if the website has .cms ending (i have my doubts on this part)
      for elem in soup.next_siblings:
        if elem.name == 'h1':
           GET THE TEXT SOME HOW
          
        break
解决方案

核心逻辑

要提取第一个<h1>之后的文本,需定位第一个<h1>元素,遍历其后续兄弟节点,收集所有文本内容直到遇到下一个<h1>节点为止。原代码的核心错误是遍历了soup的兄弟节点,而非第一个<h1>的兄弟节点。

修正后的代码

import requests
from bs4 import BeautifulSoup

url = "https://economictimes.indiatimes.com/news/international/business/volkswagen-sets-5-7-revenue-growth-target-preaches-cost-discipline/articleshow/101168014.cms"
headers = {"User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/51.0.2704.103 Safari/537.36"}

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.content, 'html.parser')

# 获取页面所有h1标签
h1_tags = soup.body.find_all('h1')
# 存在多个h1时执行提取逻辑
if len(h1_tags) > 1:
    first_h1 = h1_tags[0]
    collected_text = []
    
    # 遍历第一个h1的后续兄弟节点
    for sibling in first_h1.next_siblings:
        # 遇到下一个h1则终止遍历
        if sibling.name == 'h1':
            break
        
        # 处理文本节点
        if sibling.string:
            stripped_text = sibling.string.strip()
            if stripped_text:
                collected_text.append(stripped_text)
        # 处理标签节点,提取所有内部文本
        elif hasattr(sibling, 'get_text'):
            full_text = sibling.get_text(strip=True)
            if full_text:
                collected_text.append(full_text)
    
    # 拼接最终文本
    result = '\n'.join(collected_text)
    print(result)

关键修正点

  • 节点遍历对象修正:将soup.next_siblings改为first_h1.next_siblings,确保只遍历第一个<h1>之后的内容
  • 判断逻辑优化:将len(h1_tags) > 2改为len(h1_tags) > 1,只要存在多个<h1>就触发处理
  • 文本收集逻辑完善:同时处理纯文本节点和标签节点的文本提取,过滤空内容,避免无效文本
  • 移除冗余判断:若无需仅针对.cms后缀的网站,可直接去掉该判断;若需要,将变量i替换为当前url即可

内容的提问来源于stack exchange,提问作者Mostafa Bouzari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 07:05:41