You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python BeautifulSoup合并文章正文段落与副标题的技术问题

BeautifulSoup按原顺序提取网页副标题与正文方案

问题根因

你当前的代码是在文章容器内分别查询所有正文p标签、副标题span标签,得到两个独立的元素列表,且遍历逻辑仅处理了正文p标签,既没有拼接副标题内容,也无法保证两类内容的顺序和网页原始排版一致。

修复思路

直接遍历文章容器articulo-contenido下的所有直接子元素,逐个判断元素的类型和class属性,遇到副标题就提取副标题内容,遇到正文就提取正文内容,按遍历顺序拼接即可完全还原网页的排版顺序。

修复后代码

import urllib.request
from bs4 import BeautifulSoup

# 示例网页地址
link = "https://www.eltiempo.com/mundo/eeuu-y-canada/los-secretos-revelados-de-melania-trump-en-libro-de-stephanie-grisham-622637"
response = urllib.request.urlopen(link)
response_html = response.read().decode()
parsed_html = BeautifulSoup(response_html, 'html.parser')
page_text = ''

content_containers = parsed_html.find_all('div', class_='articulo-contenido')
for container in content_containers:
    # 遍历容器下所有直接子元素,保留原始排版顺序
    for child in container.children:
        # 过滤空白、换行等非标签节点
        if not child.name:
            continue
        # 匹配副标题
        if child.name == 'span' and 'articulo-subtitulo' in child.get('class', []):
            page_text += f"\n{child.get_text(strip=True)}"
        # 匹配正文段落
        elif child.name == 'p' and 'contenido' in child.get('class', []):
            page_text += f"\n{child.get_text(strip=True)}"

print(page_text)

可选优化

如果需要区分副标题和正文的格式,可在拼接副标题时添加强调标识,示例如下:

# 替换副标题拼接逻辑即可
page_text += f"\n*{child.get_text(strip=True)}*\n"

内容的提问来源于stack exchange,提问作者mashal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 01:54:03