You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取dmoz全层级作者分类URL漏项问题求解

问题根因

你代码漏抓链接的核心问题有两个:

  • 遍历sub_cat列表的同时直接修改列表(循环内新增子分类链接、删除当前遍历到的链接),是漏项的核心原因:Python的for循环按递增索引遍历列表,当你删除当前索引位置的元素后,列表后续元素会整体前移一位,下一轮循环索引+1时,会直接跳过刚移动到当前索引位置的元素,最终表现就是隔一个漏一个链接。
  • 硬编码切片[6:7]取第7个div.row元素的写法鲁棒性极差,一旦页面DOM结构有微小变动(比如多了一行公告、广告位),就会取错容器、漏抓链接。
优化实现方案

采用广度优先遍历(BFS)的队列逻辑分层爬取,将遍历过程和队列增删操作完全解耦,同时优化分类容器定位逻辑,避免硬编码位置:

from bs4 import BeautifulSoup
import requests
from collections import deque

# 初始化请求配置
start_path = "/Arts/Literature/Authors"
headers = {
    "User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/51.0.2704.103 Safari/537.36"
}
session = requests.Session()
session.headers.update(headers)

# 初始化待爬队列、已爬集合、结果存储
url_queue = deque([start_path])
visited = set()
all_category_paths = []

while url_queue:
    current_path = url_queue.popleft()
    # 跳过已爬路径,避免重复请求
    if current_path in visited:
        continue
    visited.add(current_path)

    current_url = f"http://dmoz.org{current_path}"
    print(f"正在爬取: {current_url}")
    resp = session.get(current_url, timeout=10)
    soup = BeautifulSoup(resp.text, "html.parser")

    # 直接通过类名定位分类面板,不再硬编码row的索引位置
    category_panel = soup.find("div", class_="panel-body")
    if not category_panel:
        # 不存在子分类面板,说明当前路径是最末级分类,存入结果
        all_category_paths.append(current_path)
        continue

    # 提取面板下所有有效分类链接
    sub_links = category_panel.find_all("a", href=True)
    has_sub_category = False
    for link in sub_links:
        href = link["href"]
        # 只保留作者分类下的站内路径,过滤外链、无关跳转链接
        if href.startswith(start_path) and href not in visited:
            url_queue.append(href)
            has_sub_category = True

    # 存在子分类的当前路径本身也是一级分类,存入结果
    if has_sub_category:
        all_category_paths.append(current_path)

# 输出结果
print(f"爬取完成,共获取{len(all_category_paths)}条分类路径:")
for path in all_category_paths:
    print(path)
方案优势
  • 队列遍历和增删操作完全解耦,从根本上解决了列表遍历过程中修改元素导致的索引错位、漏抓问题。
  • 新增visited去重集合,避免重复爬取相同路径,减少无效请求。
  • 分类容器通过panel-body类名定位,不再依赖固定索引位置,页面小幅度结构变动也不会影响抓取逻辑。
  • 增加链接过滤规则,自动排除广告、外链等无关内容,抓取结果更纯净。
  • 天然支持无限层级分类爬取,不管有多少级sub category都能完整抓取,不需要手动写多层循环。

内容的提问来源于stack exchange,提问作者Mehady

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 06:12:15