使用Python BeautifulSoup合并文章正文段落与副标题的技术问题
BeautifulSoup按原顺序提取网页副标题与正文方案
问题根因
你当前的代码是在文章容器内分别查询所有正文p标签、副标题span标签,得到两个独立的元素列表,且遍历逻辑仅处理了正文p标签,既没有拼接副标题内容,也无法保证两类内容的顺序和网页原始排版一致。
修复思路
直接遍历文章容器articulo-contenido下的所有直接子元素,逐个判断元素的类型和class属性,遇到副标题就提取副标题内容,遇到正文就提取正文内容,按遍历顺序拼接即可完全还原网页的排版顺序。
修复后代码
import urllib.request from bs4 import BeautifulSoup # 示例网页地址 link = "https://www.eltiempo.com/mundo/eeuu-y-canada/los-secretos-revelados-de-melania-trump-en-libro-de-stephanie-grisham-622637" response = urllib.request.urlopen(link) response_html = response.read().decode() parsed_html = BeautifulSoup(response_html, 'html.parser') page_text = '' content_containers = parsed_html.find_all('div', class_='articulo-contenido') for container in content_containers: # 遍历容器下所有直接子元素,保留原始排版顺序 for child in container.children: # 过滤空白、换行等非标签节点 if not child.name: continue # 匹配副标题 if child.name == 'span' and 'articulo-subtitulo' in child.get('class', []): page_text += f"\n{child.get_text(strip=True)}" # 匹配正文段落 elif child.name == 'p' and 'contenido' in child.get('class', []): page_text += f"\n{child.get_text(strip=True)}" print(page_text)
可选优化
如果需要区分副标题和正文的格式,可在拼接副标题时添加强调标识,示例如下:
# 替换副标题拼接逻辑即可 page_text += f"\n*{child.get_text(strip=True)}*\n"
内容的提问来源于stack exchange,提问作者mashal
相关产品推荐
相关产品推荐

