You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取两个Header标签间的数据并生成键值对?Header标签下文本提取方法

解决方案

一、提取Header标签下的文本

你的现有代码已经能获取Header的文本,这里补充优化写法和不同场景的提取方式:

  • 直接获取标签下所有文本(含子标签内容):item.text.strip()(加strip()去除前后空白)
  • 仅提取直接子节点的文本(排除嵌套标签内容):用' '.join(item.stripped_strings),自动过滤空白字符与换行
  • 简化现有代码的写法:
import re
from bs4 import BeautifulSoup

soup = BeautifulSoup(page.content, 'html.parser')
# 用列表推导式简化提取逻辑
htag_list = [{tag.name: tag.text.strip()} for tag in soup.find_all(re.compile('^h[1-6]'))]
print(htag_list)

二、提取两个Header之间的内容并生成键值对

核心逻辑是遍历每个Header,捕获它到下一个Header之间的所有内容,以下是两种实用实现方式:

方法1:利用find_next_siblings + 终止条件

import re
from bs4 import BeautifulSoup

soup = BeautifulSoup(page.content, 'html.parser')
headers = soup.find_all(re.compile('^h[1-6]'))
header_content_pairs = {}

for idx, header in enumerate(headers):
    # 用当前Header的文本作为键
    key = header.text.strip()
    content = []
    # 遍历当前Header之后的所有兄弟节点
    for sibling in header.find_next_siblings():
        # 遇到下一个Header就停止遍历
        if re.match('^h[1-6]$', sibling.name):
            break
        # 过滤空文本后加入内容列表
        sibling_text = sibling.text.strip()
        if sibling_text:
            content.append(sibling_text)
    # 将内容拼接成字符串存入字典
    header_content_pairs[key] = ' '.join(content)

# 输出结果
for key, val in header_content_pairs.items():
    print(f"{key}: {val}")

方法2:用next_sibling逐个遍历(精细控制)

适合处理复杂节点结构(比如内容嵌套在div中):

import re
from bs4 import BeautifulSoup

soup = BeautifulSoup(page.content, 'html.parser')
headers = soup.find_all(re.compile('^h[1-6]'))
header_content_pairs = {}

for idx, header in enumerate(headers):
    key = header.text.strip()
    content = []
    current_sibling = header.next_sibling
    while current_sibling is not None:
        # 遇到下一个Header就终止循环
        if current_sibling.name and re.match('^h[1-6]$', current_sibling.name):
            break
        # 处理标签节点的文本
        if hasattr(current_sibling, 'text'):
            text = current_sibling.text.strip()
            if text:
                content.append(text)
        # 处理纯文本节点
        else:
            str_text = str(current_sibling).strip()
            if str_text:
                content.append(str_text)
        current_sibling = current_sibling.next_sibling
    header_content_pairs[key] = ' '.join(content)

print(header_content_pairs)

注意事项

  • 如果Header仅存在于div.conWrap容器内,记得限定查找范围:soup.select_one("div.conWrap").find_all(re.compile('^h[1-6]')),避免提取页面其他区域的Header
  • 若需保留HTML格式而非纯文本,可将sibling.text替换为str(sibling),存储完整节点的HTML代码

内容的提问来源于stack exchange,提问作者Akshay Toranagatti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 18:12:12