You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup按标题爬取德语单词网站数据?

首先,报错的根源是你在遍历元素时误将NavigableString纯文本节点当成了可解析的标签节点,这类节点没有find_all方法;而去掉索引后所有分类返回相同内容,是因为代码没有限定每个分类对应的单词范围,直接全局抓取了所有单词。

解决思路

目标网站的结构是:每个h3分类标题后紧跟对应单词列表(包裹在ul标签里),分类之间以h3分隔。我们需要为每个h3限定抓取范围——只取当前h3到下一个h3之间的单词,同时过滤掉无关的文本节点。

修正后的代码示例

import requests
from bs4 import BeautifulSoup

url = "https://almancakonulari.com/a1-seviye-almanca-kelimeler/#gsc.tab=0"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")

# 获取所有分类标题h3
categories = soup.find_all("h3")
german_dict = {}

for category in categories:
    category_name = category.get_text(strip=True)
    german_dict[category_name] = []
    
    # 从当前h3的下一个兄弟元素开始遍历
    next_elem = category.next_sibling
    while next_elem is not None:
        # 过滤纯文本节点(比如换行、空格)
        if next_elem.name == "ul":
            # 提取ul下所有li的单词内容
            words = [li.get_text(strip=True) for li in next_elem.find_all("li")]
            german_dict[category_name].extend(words)
        # 遇到下一个h3就停止当前分类的抓取
        elif next_elem.name == "h3":
            break
        # 继续遍历下一个兄弟元素
        next_elem = next_elem.next_sibling

# 打印示例结果
for cat, words in german_dict.items():
    print(f"【{cat}】")
    print(words[:5])  # 仅显示前5个单词

关键说明

  1. 过滤文本节点:通过判断next_elem.name是否存在,排除换行、空格这类NavigableString节点,避免调用find_all时报错。
  2. 限定抓取范围:遍历兄弟元素时,一旦遇到下一个h3就终止当前分类的单词收集,确保每个分类只对应自身的单词列表。
  3. 精准定位容器:网站中每个分类的单词都放在h3后的ul标签里,直接定位ul再提取li内容,比全局查找更高效准确。

替代方案(用find_next_siblings)

如果觉得循环遍历麻烦,也可以用find_next_siblings方法截取到下一个h3之前的元素:

for category in categories:
    category_name = category.get_text(strip=True)
    # 获取当前h3之后的所有兄弟元素
    siblings = category.find_next_siblings()
    # 找到下一个h3的索引,截取之前的元素
    next_h3_idx = None
    for i, sib in enumerate(siblings):
        if sib.name == "h3":
            next_h3_idx = i
            break
    target_siblings = siblings[:next_h3_idx] if next_h3_idx else siblings
    
    # 提取目标元素中的所有单词
    words = []
    for sib in target_siblings:
        if sib.name == "ul":
            words.extend([li.get_text(strip=True) for li in sib.find_all("li")])
    german_dict[category_name] = words

内容的提问来源于stack exchange,提问作者NewPartizal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 07:45:36