如何提取第一个h1标签后的文本?网页批量爬取技术方案
问题
每日需爬取并清洗100个网站文本,遇到一类存在多个<h1>标签的网站:滚动至下一个<h1>标签时URL会变化(示例地址:https://economictimes.indiatimes.com/news/international/business/volkswagen-sets-5-7-revenue-growth-target-preaches-cost-discipline/articleshow/101168014.cms)。需提取第一个<h1>标签之后的文本(该文本不在<p>标签内),现有代码逻辑存在问题,请求修正。
现有代码:
response=requests.get('https://economictimes.indiatimes.com/news/international/business/volkswagen-sets-5-7-revenue-growth-target-preaches-cost-discipline/articleshow/101168014.cms',headers={"User-Agent" : "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/51.0.2704.103 Safari/537.36"}) soup = BeautifulSoup(response.content, 'html.parser') if len(soup.body.find_all('h1'))>2: #to check if there is more than one tag if i.endswith(".cms"): #to check if the website has .cms ending (i have my doubts on this part) for elem in soup.next_siblings: if elem.name == 'h1': GET THE TEXT SOME HOW break
解决方案
核心逻辑
要提取第一个<h1>之后的文本,需定位第一个<h1>元素,遍历其后续兄弟节点,收集所有文本内容直到遇到下一个<h1>节点为止。原代码的核心错误是遍历了soup的兄弟节点,而非第一个<h1>的兄弟节点。
修正后的代码
import requests from bs4 import BeautifulSoup url = "https://economictimes.indiatimes.com/news/international/business/volkswagen-sets-5-7-revenue-growth-target-preaches-cost-discipline/articleshow/101168014.cms" headers = {"User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/51.0.2704.103 Safari/537.36"} response = requests.get(url, headers=headers) soup = BeautifulSoup(response.content, 'html.parser') # 获取页面所有h1标签 h1_tags = soup.body.find_all('h1') # 存在多个h1时执行提取逻辑 if len(h1_tags) > 1: first_h1 = h1_tags[0] collected_text = [] # 遍历第一个h1的后续兄弟节点 for sibling in first_h1.next_siblings: # 遇到下一个h1则终止遍历 if sibling.name == 'h1': break # 处理文本节点 if sibling.string: stripped_text = sibling.string.strip() if stripped_text: collected_text.append(stripped_text) # 处理标签节点,提取所有内部文本 elif hasattr(sibling, 'get_text'): full_text = sibling.get_text(strip=True) if full_text: collected_text.append(full_text) # 拼接最终文本 result = '\n'.join(collected_text) print(result)
关键修正点
- 节点遍历对象修正:将
soup.next_siblings改为first_h1.next_siblings,确保只遍历第一个<h1>之后的内容 - 判断逻辑优化:将
len(h1_tags) > 2改为len(h1_tags) > 1,只要存在多个<h1>就触发处理 - 文本收集逻辑完善:同时处理纯文本节点和标签节点的文本提取,过滤空内容,避免无效文本
- 移除冗余判断:若无需仅针对
.cms后缀的网站,可直接去掉该判断;若需要,将变量i替换为当前url即可
内容的提问来源于stack exchange,提问作者Mostafa Bouzari
相关产品推荐
相关产品推荐

