如何用Python+BeautifulSoup爬取文章并保留原网页h2标签格式?
爬取文章时段落文本重复输出的解决方法
待爬取的HTML结构
<section name="articleBody"> <h2 id="link-a1b2c3d4" style="color:rgb(18, 18, 18)">Point 1</h2> <p>Text on Point 1</p> <h2 id="link-a1b2c3d4" style="color:rgb(18, 18, 18)">Point 2</h2> <p>Text on Point 2</p> <h2 id="link-a1b2c3d4" style="color:rgb(18, 18, 18)">Point 3</h2> <p>Text on Point 3</p> </section>
期望输出
<h2> Point 1 </h2> Text on Point 1 <h2> Point 2 </h2> Text on Point 2 <h2> Point 3 </h2> Text on Point 3
尝试的代码
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import re import requests import os url = "http://www.test.com" options = Options() options.headless = True driver = webdriver.Chrome(options=options) driver.get(url) # Parse HTML soup = BeautifulSoup(driver.page_source, "html.parser") # Extract article text and include image tags article_body = soup.find("section", {"name": "articleBody"}) article_text = "" # Iterate through each element in article_body def process_elements(elements): global article_text for element in elements: if element.name == "div": # Check for div elements if "h2" in element.find_all(True): # Check if div contains h2 article_text += f"<h2>{element.find('h2').get_text()}</h2>\n" else: article_text += element.get_text(strip=True) + "\n" process_elements(element.children) # Recurse through nested elements else: # Handle non-div elements directly if element.name == "h2": article_text += f"<h2>{element.get_text()}</h2>\n" else: article_text += element.get_text(strip=True) + "\n" # Apply the processing function to article_body and its children process_elements(article_body.children) print(article_text)
问题原因
重复输出的核心原因是重复处理元素:
- 处理div时,直接提取了div的文本内容(包含子元素的文本)
- 随后又递归调用
process_elements处理div的子元素,导致子元素的文本被再次添加到结果中
解决方法
修正思路
- 移除全局变量,改用函数返回字符串的方式,避免全局状态混乱
- 处理div时,不直接提取文本,仅递归处理其子元素
- 跳过空白的文本节点,避免多余空行
- 仅在文本内容非空时添加换行,优化输出格式
修正后的代码
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options url = "http://www.test.com" options = Options() options.headless = True driver = webdriver.Chrome(options=options) driver.get(url) # Parse HTML soup = BeautifulSoup(driver.page_source, "html.parser") # 定位文章主体 article_body = soup.find("section", {"name": "articleBody"}) def process_elements(elements): result = "" for element in elements: # 跳过空白的文本节点 if element.name is None: stripped_text = element.strip() if stripped_text: result += stripped_text + "\n" continue if element.name == "div": # 递归处理div的子元素,不直接提取div文本 result += process_elements(element.children) else: if element.name == "h2": result += f"<h2> {element.get_text(strip=True)} </h2>\n" else: # 处理其他元素,仅保留非空文本 text = element.get_text(strip=True) if text: result += text + "\n" return result # 生成最终文本并输出 article_text = process_elements(article_body.children) print(article_text.strip())
关键修正点说明
- 避免重复处理:div元素不再直接提取文本,仅递归处理子元素,确保每个元素只被处理一次
- 全局变量替换:函数通过返回值拼接字符串,避免全局变量带来的状态冲突
- 空白节点过滤:跳过无意义的空白文本节点,减少多余空行
- 非空文本判断:仅当提取到的文本内容非空时才添加换行,输出更整洁
内容的提问来源于stack exchange,提问作者PandaBlink
相关产品推荐
相关产品推荐

