You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+BeautifulSoup爬取文章并保留原网页h2标签格式?

爬取文章时段落文本重复输出的解决方法

待爬取的HTML结构

<section name="articleBody">
  <h2 id="link-a1b2c3d4" style="color:rgb(18, 18, 18)">Point 1</h2>
  <p>Text on Point 1</p>
  <h2 id="link-a1b2c3d4" style="color:rgb(18, 18, 18)">Point 2</h2>
  <p>Text on Point 2</p>
  <h2 id="link-a1b2c3d4" style="color:rgb(18, 18, 18)">Point 3</h2>
  <p>Text on Point 3</p>
</section>

期望输出

<h2> Point 1 </h2>
Text on Point 1
<h2> Point 2 </h2>
Text on Point 2
<h2> Point 3 </h2>
Text on Point 3

尝试的代码

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import re
import requests
import os

url = "http://www.test.com"
options = Options()
options.headless = True
driver = webdriver.Chrome(options=options)
driver.get(url)

# Parse HTML
soup = BeautifulSoup(driver.page_source, "html.parser")

# Extract article text and include image tags
article_body = soup.find("section", {"name": "articleBody"})
article_text = ""

# Iterate through each element in article_body
def process_elements(elements):
    global article_text
    for element in elements:
        if element.name == "div":  # Check for div elements
            if "h2" in element.find_all(True):  # Check if div contains h2
                article_text += f"<h2>{element.find('h2').get_text()}</h2>\n"
            else:
                article_text += element.get_text(strip=True) + "\n"
            process_elements(element.children)  # Recurse through nested elements
        else:
            # Handle non-div elements directly
            if element.name == "h2":
                article_text += f"<h2>{element.get_text()}</h2>\n"
            else:
                article_text += element.get_text(strip=True) + "\n"


# Apply the processing function to article_body and its children
process_elements(article_body.children)

print(article_text)

问题原因

重复输出的核心原因是重复处理元素:

  1. 处理div时,直接提取了div的文本内容(包含子元素的文本)
  2. 随后又递归调用process_elements处理div的子元素,导致子元素的文本被再次添加到结果中

解决方法

修正思路

  • 移除全局变量,改用函数返回字符串的方式,避免全局状态混乱
  • 处理div时,不直接提取文本,仅递归处理其子元素
  • 跳过空白的文本节点,避免多余空行
  • 仅在文本内容非空时添加换行,优化输出格式

修正后的代码

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options

url = "http://www.test.com"
options = Options()
options.headless = True
driver = webdriver.Chrome(options=options)
driver.get(url)

# Parse HTML
soup = BeautifulSoup(driver.page_source, "html.parser")

# 定位文章主体
article_body = soup.find("section", {"name": "articleBody"})

def process_elements(elements):
    result = ""
    for element in elements:
        # 跳过空白的文本节点
        if element.name is None:
            stripped_text = element.strip()
            if stripped_text:
                result += stripped_text + "\n"
            continue
        
        if element.name == "div":
            # 递归处理div的子元素,不直接提取div文本
            result += process_elements(element.children)
        else:
            if element.name == "h2":
                result += f"<h2> {element.get_text(strip=True)} </h2>\n"
            else:
                # 处理其他元素,仅保留非空文本
                text = element.get_text(strip=True)
                if text:
                    result += text + "\n"
    return result

# 生成最终文本并输出
article_text = process_elements(article_body.children)
print(article_text.strip())

关键修正点说明

  1. 避免重复处理:div元素不再直接提取文本,仅递归处理子元素,确保每个元素只被处理一次
  2. 全局变量替换:函数通过返回值拼接字符串,避免全局变量带来的状态冲突
  3. 空白节点过滤:跳过无意义的空白文本节点,减少多余空行
  4. 非空文本判断:仅当提取到的文本内容非空时才添加换行,输出更整洁

内容的提问来源于stack exchange,提问作者PandaBlink

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 05:55:54