You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何以更Pythonic的方式捕获多个同类型标签,实现按H3分类存储下载链接?

解决爬虫捕获单个键值对的问题并优化实现

你的思路非常清晰,但原代码的逻辑存在一个关键漏洞:只有当遇到下一个<h3>标签时,才会将当前标题和链接存入字典,这就导致最后一组标题对应的链接永远无法被添加,所以你只能拿到一个键值对。下面是更Pythonic的优化实现,同时修复了这个问题:

优化后的代码

import requests
from bs4 import BeautifulSoup

def scrape():
    resp = requests.get('https://www.animeout.xyz/love-live-nijigasaki-gakuen-school-idol-doukoukai-1080p-300mb720p-150mbepisode-1/')
    soup = BeautifulSoup(resp.text, 'html.parser')
    content_div = soup.find('div', class_='article-content')
    output = {}
    current_title = None

    # 只遍历目标div下的直接h3和p标签,过滤无关节点
    for tag in content_div.find_all(['h3', 'p'], recursive=False):
        if tag.name == 'h3':
            # 初始化当前标题对应的空列表
            current_title = tag.get_text(strip=True)
            output[current_title] = []
        elif tag.name == 'p' and current_title is not None:
            # 提取p标签内的所有Direct Download链接
            direct_links = tag.find_all('a', string='Direct Download')
            # 用生成器表达式简化链接收集
            output[current_title].extend(link.get('href') for link in direct_links if link.get('href'))
    
    return output

# 测试运行
if __name__ == '__main__':
    download_data = scrape()
    for title, links in download_data.items():
        print(f"*{title}*")
        for idx, link in enumerate(links, 1):
            print(f"  {idx}. `{link}`")

优化点说明

  • 逻辑更清晰:用current_title追踪当前对应的标题,遇到<h3>就初始化字典键和空列表,遇到<p>直接将链接追加到对应列表,彻底避免了遗漏最后一组数据的问题。
  • 过滤无关节点:通过find_all(['h3', 'p'], recursive=False)只筛选目标div的直接子标签中的h3和p,跳过了多余的文本节点(比如换行、空格),无需处理sibling的复杂判断。
  • 更健壮的链接提取:用extend+生成器表达式简化链接收集,同时判断link.get('href')不为空,避免存入无效链接。
  • 干净的标题文本:get_text(strip=True)自动去除标题中的多余空格、换行和制表符,让字典键更整洁。

内容的提问来源于stack exchange,提问作者Hans

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 10:12:31