网站子板块URL爬取异常:无法获取全部子页面链接
问题与解决方案
问题说明
我尝试爬取「Estadístiques」板块下的所有子板块URL并生成列表,原以为代码可正常运行,但发现例如「Estadística de l’ensenyament 2021-2022」这类板块的子页面链接未被全部抓取。
原代码
import requests from bs4 import BeautifulSoup import pandas as pd education_statistic_section = "https://educacio.gencat.cat/ca/departament/estadistiques/" # 主页面,包含教育统计板块的子板块 -- 0级页面 html_eduaction_section_levels = "distribuidora-item grey" # 包含各教育板块层级的类名 web_education = "https://educacio.gencat.cat" list_title_level_1 = [] list_web_level_1 = [] list_all_web=[] list_title_subsecction = [] list_web_subsecction = [] list_title_document = [] list_web_document = [] def parse_url(url): response = requests.get(url) content = response.content parsed_response = BeautifulSoup(content, "lxml") return parsed_response def first_secction_statistic (): # 抓取主页面中所有class为distribuidora-item grey的div的标题和链接 soup = parse_url(education_statistic_section) # 解析主页面 html_div_level_1 = soup.find_all('div', {'class':html_eduaction_section_levels}) # 获取所有目标div元素 for html_elements_level_1 in html_div_level_1: # 遍历每个div元素 list_title_level_1.append( html_elements_level_1.text.strip()) # 提取标题并加入列表 html_tags_as_level_1= html_elements_level_1.find('a') # 获取a标签 list_web_level_1.append(web_education+html_tags_as_level_1.get('href'))# 拼接完整URL并加入列表 first_secction_statistic() # csv = pd.DataFrame({'Títols nivell 1': pd.Series(list_title_level_1), 'Web nivell 1': pd.Series(list_web_level_1)}) def all_web_subsecction_statistic (): # 列出教育板块的所有子板块链接 for i in list_web_level_1: # 遍历一级页面列表 soup = None # 清空变量 soup = parse_url(i) # 解析当前页面 html_tags_a = soup.find_all('a') # 获取所有a标签(包含文档或相关链接) for element in html_tags_a: # 遍历每个a标签 str_element = str(element.get('href')) # 获取href属性并转为字符串 if str_element.startswith('/ca/departament/estadistiques/'): # 判断是否为统计板块的子链接 subsecction_web_statistic= web_education+str_element # 拼接完整URL if subsecction_web_statistic not in list_all_web: # 避免重复添加 list_all_web.append(subsecction_web_statistic) all_web_subsecction_statistic() print(list_all_web)
问题原因
原代码仅处理了**一级页面(list_web_level_1)**中的链接,没有递归爬取这些子页面中更深层级的符合条件的链接。比如「Estadística de l’ensenyament 2021-2022」页面里的子板块链接,代码没有对这些新抓取到的链接再次进行解析,导致遗漏深层子页面。
修改方案
改用队列+循环的方式,持续处理待爬取的链接,直到所有符合条件的子链接都被抓取。这种方式能自动处理多层级的页面结构,不会遗漏深层链接。
修改后的代码
import requests from bs4 import BeautifulSoup import pandas as pd education_statistic_section = "https://educacio.gencat.cat/ca/departament/estadistiques/" html_eduaction_section_levels = "distribuidora-item grey" web_education = "https://educacio.gencat.cat" list_title_level_1 = [] list_web_level_1 = [] list_all_web = [] visited_urls = set() # 用集合存储已访问的URL,避免重复爬取 def parse_url(url): try: response = requests.get(url) response.raise_for_status() # 捕获HTTP请求错误 return BeautifulSoup(response.content, "lxml") except requests.exceptions.RequestException as e: print(f"请求URL失败: {url}, 错误信息: {e}") return None def first_secction_statistic(): soup = parse_url(education_statistic_section) if not soup: return html_div_level_1 = soup.find_all('div', {'class': html_eduaction_section_levels}) for html_elements_level_1 in html_div_level_1: list_title_level_1.append(html_elements_level_1.text.strip()) html_tags_as_level_1 = html_elements_level_1.find('a') if html_tags_as_level_1 and html_tags_as_level_1.get('href'): full_url = web_education + html_tags_as_level_1.get('href') list_web_level_1.append(full_url) visited_urls.add(full_url) # 标记为已访问 list_all_web.append(full_url) # 加入总列表 def crawl_all_subsections(): # 初始化待爬取队列,先加入一级页面 crawl_queue = list_web_level_1.copy() while crawl_queue: current_url = crawl_queue.pop(0) # 取出队列第一个URL soup = parse_url(current_url) if not soup: continue html_tags_a = soup.find_all('a') for element in html_tags_a: href = element.get('href') if not href: continue str_element = str(href) # 判断是否为统计板块的子链接,且不是文档类链接(比如.pdf等) if str_element.startswith('/ca/departament/estadistiques/') and not str_element.endswith(('.pdf', '.xls', '.xlsx', '.doc', '.docx')): full_url = web_education + str_element if full_url not in visited_urls: visited_urls.add(full_url) list_all_web.append(full_url) crawl_queue.append(full_url) # 将新发现的链接加入队列,等待爬取 # 执行抓取 first_secction_statistic() crawl_all_subsections() # 输出结果 print("所有子板块URL列表:") for url in list_all_web: print(url) # 可选:保存为CSV # pd.DataFrame({'所有子板块URL': list_all_web}).to_csv('教育统计子板块URL.csv', index=False, encoding='utf-8-sig')
修改点说明
- 新增
visited_urls集合:用于记录已爬取的URL,彻底避免重复爬取和循环爬取。 - 改用队列
crawl_queue:通过循环持续处理队列中的URL,每次爬取到新的符合条件的链接就加入队列,实现递归式的多层级爬取。 - 增加请求异常处理:捕获HTTP请求错误,避免单个URL请求失败导致整个程序崩溃。
- 过滤文档类链接:新增对
.pdf、.xls等文档后缀的判断,避免把文档下载链接误当成子板块链接。
内容的提问来源于stack exchange,提问作者Merinoide
相关产品推荐
相关产品推荐

