You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup获取h3标签内的指定内容?

解决德语词典分类提取与嵌套JSON构建问题

需求说明

需要提取HTML中h3标签内两个span之间的文本(如示例中的"Zahlen"),并将后续表格中的单词按该分类组织成嵌套JSON结构,替代当前的平级单词列表。

目标HTML结构

<h3>
  <span class="ez-toc-section" id="Zahlen"></span>
     Zahlen
  <span class="ez-toc-section-end"></span>
</h3>

现有代码(无法按分类分组)

result = requests.get(url)
soup = BeautifulSoup(result.text, 'html.parser')
div = soup.find('div', class_='entry-content clearfix')


for word in div.find_all('table', class_='table table-bordered'):
    for word1 in word.find_all('tbody'):
        rows = word1.find_all('tr')
        for row in rows:
            each_word = row.find_all('td')
            case = {
                "index": each_word[0].string,
                "word": each_word[1].string,
                "meaning": each_word[2].string
            }
            list.append(case)

with open('DictionaryWords.json', 'w', encoding='utf-8') as f:
    json.dump(list, f, ensure_ascii=False, indent=4)

当前输出(平级列表)

[
    {
        "index": "1.",
        "word": "Hallo",
        "meaning": "Merhaba"
    },
    {
        "index": "2.",
        "word": "Herzlich willkommen",
        "meaning": "Hoş geldiniz"
    },
    {
        "index": "3.",
        "word": "Auf Wiedersehen",
        "meaning": "Hoşça kalın"
    },
    {
        "index": "4.",
        "word": "Guten Morgen",
        "meaning": "Günaydın"
    },
    {
        "index": "5.",
        "word": "Haben Sie einen guten Tag",
        "meaning": "İyi günler"
    }
]

期望输出(分类嵌套结构)

[
     {
      "zahlen":[
        {
            "index": "1.",
            "word": "Hallo",
            "meaning": "Merhaba"
        },
        {
            "index": "2.",
            "word": "Herzlich willkommen",
            "meaning": "Hoş geldiniz"
        },
        {
            "index": "3.",
            "word": "Auf Wiedersehen",
            "meaning": "Hoşça kalın"
        }
      ]
     }
]

修改后的代码

import requests
from bs4 import BeautifulSoup
import json

result = requests.get(url)
soup = BeautifulSoup(result.text, 'html.parser')
div = soup.find('div', class_='entry-content clearfix')

# 存储最终分类结构
dictionary = []
current_category = None
current_words = []

# 按顺序遍历div下的子元素,关联h3与对应表格
for element in div.children:
    # 处理h3标签,更新当前分类
    if element.name == 'h3':
        # 先保存上一个分类的内容(如果存在)
        if current_category and current_words:
            dictionary.append({current_category.lower(): current_words})
            current_words = []
        # 提取h3纯文本并清理空白
        current_category = element.get_text(strip=True)
    # 处理表格,将单词添加到当前分类列表
    elif element.name == 'table' and 'table-bordered' in element.get('class', []):
        if current_category:
            tbody = element.find('tbody')
            if tbody:
                rows = tbody.find_all('tr')
                for row in rows:
                    each_word = row.find_all('td')
                    # 确保表格列数足够,避免索引错误
                    if len(each_word) >= 3:
                        case = {
                            "index": each_word[0].string.strip() if each_word[0].string else "",
                            "word": each_word[1].string.strip() if each_word[1].string else "",
                            "meaning": each_word[2].string.strip() if each_word[2].string else ""
                        }
                        current_words.append(case)

# 处理最后一个未保存的分类
if current_category and current_words:
    dictionary.append({current_category.lower(): current_words})

# 写入JSON文件
with open('DictionaryWords.json', 'w', encoding='utf-8') as f:
    json.dump(dictionary, f, ensure_ascii=False, indent=4)

关键修改说明

  1. 按顺序关联分类与表格:不再批量提取所有表格,而是遍历div的子元素,将每个表格与最近的h3分类绑定
  2. 提取h3纯文本:使用get_text(strip=True)自动忽略span标签,直接提取并清理中间的分类文本
  3. 空值与异常处理:增加列数判断和空字符串处理,避免因表格内容缺失导致的报错
  4. 分类名格式统一:将分类名转为小写,匹配期望输出的键名格式

内容的提问来源于stack exchange,提问作者NewPartizal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 07:25:14