You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用BeautifulSoup抓取h3/h4标签时缺失对应标签如何返回空值

问题原因

你现有的实现逻辑是分别抓取所有<h3>和<h4>节点生成两个独立列表,再用zip_longest直接按索引对齐配对,完全没有考虑两类标签在页面中穿插出现的原始顺序,自然会出现配对错位的问题。

实现思路

要按页面实际出现顺序配对,遵循以下逻辑即可:

  • 一次性提取页面中所有<h3>和<h4>节点,保留原始先后顺序
  • 遍历节点列表,每遇到一个<h4>就向后寻找第一个未被匹配的<h3>配对,找不到对应<h3>时填充空值
  • 已配对的<h3>标记为已使用,不再参与后续配对
完整实现代码
import requests
from bs4 import BeautifulSoup

url = requests.get('https://example.com')
soup = BeautifulSoup(url.text, 'lxml')

# 一次性提取所有h3、h4节点,保留页面原始顺序
all_nodes = soup.find_all(['h3', 'h4'])
result = []
used_h3 = set()

for i, node in enumerate(all_nodes):
    if node.name == 'h4':
        h4_text = node.text.strip()
        h3_text = ''
        # 向后查找第一个未被使用的h3
        for j in range(i+1, len(all_nodes)):
            next_node = all_nodes[j]
            if next_node.name == 'h3' and id(next_node) not in used_h3:
                h3_text = next_node.text.strip()
                used_h3.add(id(next_node))
                break
        result.append([h4_text, h3_text])

# 输出配对结果
for row in result:
    print(row)
效果验证

对应你给出的示例页面结构,运行以上代码输出结果为:

['xxx', 'yyy'], ['sss', ''], ['zzz', 'ooo']

完全匹配预期结果。

内容的提问来源于stack exchange,提问作者user14878631

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 05:54:04