You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

同结构学区校网爬虫仅单站生效,换站输出空白求排查

爬虫适配不同域名网站的问题排查与修复

问题背景

爬取某学区学校网站,网站结构一致仅域名前缀不同,但爬虫在Fairfax High School网站正常运行,切换到Aldrin School域名后输出空白。

原代码

import pandas as pd
from bs4 import BeautifulSoup
import requests

url = 'https://fairfaxhs.fcps.edu/staff-directory?field_last_name_from=&field_last_name_to=&items_per_page=10&keywords=&page={page}'
data = []
for page in range(0,2):
    soup = BeautifulSoup(requests.get(url.format(page=page)).text,'lxml')
    try:
        for u in ['https://fairfaxhs.fcps.edu'+link.a.get('href') for link in soup.table.select('tr td[class="views-field views-field-rendered-item"]')]:
            soup2 = BeautifulSoup(requests.get(u).text,'lxml')
            d={
                'Name': soup2.select_one('h1.node__title.fcps-color--dark11').get_text(strip=True), 
                'Position': soup2.select_one('h1+div').get_text(strip=True),
                'contact_url': u
                }
            data.append(d)
    except:
        pass

df=pd.DataFrame(data).to_csv('fcps_school.csv',index=False)
print(df)

目标爬取URL

https://aldrines.fcps.edu/staff-directory?keywords=&field_last_name_from=&field_last_name_to=&items_per_page=10&page=
https://aldrines.fcps.edu

问题排查与修复

核心问题点

  1. 硬编码域名未全局替换:原代码中构造详情页URL时,硬写了https://fairfaxhs.fcps.edu前缀,切换域名时仅修改了列表页URL,导致详情页请求的还是旧域名地址,自然无法获取数据。
  2. 异常捕获过于宽泛:except: pass会吞掉所有错误(比如请求失败、元素找不到等),无法定位具体问题,调试难度大。

修改后的代码

import pandas as pd
from bs4 import BeautifulSoup
import requests

# 只需修改此处域名前缀即可切换目标学校
school_domain = "aldrines.fcps.edu"
base_url = f"https://{school_domain}"
staff_list_url = f"{base_url}/staff-directory?field_last_name_from=&field_last_name_to=&items_per_page=10&keywords=&page={{page}}"

data = []
for page in range(0, 2):
    try:
        # 请求列表页并检查HTTP状态
        list_res = requests.get(staff_list_url.format(page=page))
        list_res.raise_for_status()
        soup = BeautifulSoup(list_res.text, 'lxml')
        
        # 提取员工详情页链接
        staff_cells = soup.table.select('tr td[class="views-field views-field-rendered-item"]')
        for cell in staff_cells:
            try:
                detail_href = cell.a.get('href')
                detail_url = f"{base_url}{detail_href}"
                
                # 请求详情页并检查状态
                detail_res = requests.get(detail_url)
                detail_res.raise_for_status()
                soup2 = BeautifulSoup(detail_res.text, 'lxml')
                
                # 提取数据
                name = soup2.select_one('h1.node__title.fcps-color--dark11').get_text(strip=True)
                position = soup2.select_one('h1+div').get_text(strip=True)
                
                data.append({
                    'Name': name,
                    'Position': position,
                    'contact_url': detail_url
                })
            except Exception as e:
                print(f"处理详情页失败 {detail_url}: {str(e)}")
                continue
    except Exception as e:
        print(f"请求列表页(页码{page})失败: {str(e)}")
        continue

# 保存并输出结果
df = pd.DataFrame(data)
df.to_csv('fcps_school.csv', index=False)
print(df)

关键优化点

  • 用school_domain变量统一管理域名,切换目标学校时仅需修改一行代码
  • 增加raise_for_status()检查HTTP请求是否成功,避免请求失败后继续无效解析
  • 拆分异常捕获逻辑,分别处理列表页和详情页的错误,并打印具体信息,便于调试
  • 移除宽泛的except: pass,精准定位问题

内容的提问来源于stack exchange,提问作者Syed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 23:35:41