You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网站子板块URL爬取异常:无法获取全部子页面链接

问题与解决方案

问题说明

我尝试爬取「Estadístiques」板块下的所有子板块URL并生成列表,原以为代码可正常运行,但发现例如「Estadística de l’ensenyament 2021-2022」这类板块的子页面链接未被全部抓取。

原代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

education_statistic_section = "https://educacio.gencat.cat/ca/departament/estadistiques/" # 主页面,包含教育统计板块的子板块 -- 0级页面
html_eduaction_section_levels = "distribuidora-item grey" # 包含各教育板块层级的类名
web_education = "https://educacio.gencat.cat"

list_title_level_1 = []
list_web_level_1 = []
list_all_web=[]

list_title_subsecction = []
list_web_subsecction = []
list_title_document = []
list_web_document = []

def parse_url(url):
    response = requests.get(url)
    content = response.content
    parsed_response = BeautifulSoup(content, "lxml")
    return parsed_response

def first_secction_statistic (): # 抓取主页面中所有class为distribuidora-item grey的div的标题和链接
    soup = parse_url(education_statistic_section) # 解析主页面
    html_div_level_1 = soup.find_all('div', {'class':html_eduaction_section_levels}) # 获取所有目标div元素
    for html_elements_level_1 in html_div_level_1: # 遍历每个div元素
        list_title_level_1.append( html_elements_level_1.text.strip()) # 提取标题并加入列表
        html_tags_as_level_1= html_elements_level_1.find('a') # 获取a标签
        list_web_level_1.append(web_education+html_tags_as_level_1.get('href'))# 拼接完整URL并加入列表

first_secction_statistic()
# csv = pd.DataFrame({'Títols nivell 1': pd.Series(list_title_level_1), 'Web nivell 1': pd.Series(list_web_level_1)})

def all_web_subsecction_statistic (): # 列出教育板块的所有子板块链接
    for i in list_web_level_1: # 遍历一级页面列表
        soup = None # 清空变量
        soup = parse_url(i) # 解析当前页面
        html_tags_a = soup.find_all('a') # 获取所有a标签(包含文档或相关链接)
        for element in html_tags_a: # 遍历每个a标签
            str_element = str(element.get('href')) # 获取href属性并转为字符串
            if str_element.startswith('/ca/departament/estadistiques/'): # 判断是否为统计板块的子链接
                subsecction_web_statistic= web_education+str_element # 拼接完整URL
                if subsecction_web_statistic not in list_all_web: # 避免重复添加
                    list_all_web.append(subsecction_web_statistic)

all_web_subsecction_statistic()
print(list_all_web)

问题原因

原代码仅处理了**一级页面(list_web_level_1)**中的链接,没有递归爬取这些子页面中更深层级的符合条件的链接。比如「Estadística de l’ensenyament 2021-2022」页面里的子板块链接,代码没有对这些新抓取到的链接再次进行解析,导致遗漏深层子页面。

修改方案

改用队列+循环的方式,持续处理待爬取的链接,直到所有符合条件的子链接都被抓取。这种方式能自动处理多层级的页面结构,不会遗漏深层链接。

修改后的代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

education_statistic_section = "https://educacio.gencat.cat/ca/departament/estadistiques/"
html_eduaction_section_levels = "distribuidora-item grey"
web_education = "https://educacio.gencat.cat"

list_title_level_1 = []
list_web_level_1 = []
list_all_web = []
visited_urls = set() # 用集合存储已访问的URL,避免重复爬取

def parse_url(url):
    try:
        response = requests.get(url)
        response.raise_for_status() # 捕获HTTP请求错误
        return BeautifulSoup(response.content, "lxml")
    except requests.exceptions.RequestException as e:
        print(f"请求URL失败: {url}, 错误信息: {e}")
        return None

def first_secction_statistic():
    soup = parse_url(education_statistic_section)
    if not soup:
        return
    html_div_level_1 = soup.find_all('div', {'class': html_eduaction_section_levels})
    for html_elements_level_1 in html_div_level_1:
        list_title_level_1.append(html_elements_level_1.text.strip())
        html_tags_as_level_1 = html_elements_level_1.find('a')
        if html_tags_as_level_1 and html_tags_as_level_1.get('href'):
            full_url = web_education + html_tags_as_level_1.get('href')
            list_web_level_1.append(full_url)
            visited_urls.add(full_url) # 标记为已访问
            list_all_web.append(full_url) # 加入总列表

def crawl_all_subsections():
    # 初始化待爬取队列,先加入一级页面
    crawl_queue = list_web_level_1.copy()
    
    while crawl_queue:
        current_url = crawl_queue.pop(0) # 取出队列第一个URL
        soup = parse_url(current_url)
        if not soup:
            continue
        
        html_tags_a = soup.find_all('a')
        for element in html_tags_a:
            href = element.get('href')
            if not href:
                continue
            str_element = str(href)
            # 判断是否为统计板块的子链接,且不是文档类链接(比如.pdf等)
            if str_element.startswith('/ca/departament/estadistiques/') and not str_element.endswith(('.pdf', '.xls', '.xlsx', '.doc', '.docx')):
                full_url = web_education + str_element
                if full_url not in visited_urls:
                    visited_urls.add(full_url)
                    list_all_web.append(full_url)
                    crawl_queue.append(full_url) # 将新发现的链接加入队列,等待爬取

# 执行抓取
first_secction_statistic()
crawl_all_subsections()

# 输出结果
print("所有子板块URL列表:")
for url in list_all_web:
    print(url)

# 可选:保存为CSV
# pd.DataFrame({'所有子板块URL': list_all_web}).to_csv('教育统计子板块URL.csv', index=False, encoding='utf-8-sig')

修改点说明

  1. 新增visited_urls集合:用于记录已爬取的URL,彻底避免重复爬取和循环爬取。
  2. 改用队列crawl_queue:通过循环持续处理队列中的URL,每次爬取到新的符合条件的链接就加入队列,实现递归式的多层级爬取。
  3. 增加请求异常处理:捕获HTTP请求错误,避免单个URL请求失败导致整个程序崩溃。
  4. 过滤文档类链接:新增对.pdf、.xls等文档后缀的判断,避免把文档下载链接误当成子板块链接。

内容的提问来源于stack exchange,提问作者Merinoide

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 10:20:31