同结构学区校网爬虫仅单站生效,换站输出空白求排查
爬虫适配不同域名网站的问题排查与修复
问题背景
爬取某学区学校网站,网站结构一致仅域名前缀不同,但爬虫在Fairfax High School网站正常运行,切换到Aldrin School域名后输出空白。
原代码
import pandas as pd from bs4 import BeautifulSoup import requests url = 'https://fairfaxhs.fcps.edu/staff-directory?field_last_name_from=&field_last_name_to=&items_per_page=10&keywords=&page={page}' data = [] for page in range(0,2): soup = BeautifulSoup(requests.get(url.format(page=page)).text,'lxml') try: for u in ['https://fairfaxhs.fcps.edu'+link.a.get('href') for link in soup.table.select('tr td[class="views-field views-field-rendered-item"]')]: soup2 = BeautifulSoup(requests.get(u).text,'lxml') d={ 'Name': soup2.select_one('h1.node__title.fcps-color--dark11').get_text(strip=True), 'Position': soup2.select_one('h1+div').get_text(strip=True), 'contact_url': u } data.append(d) except: pass df=pd.DataFrame(data).to_csv('fcps_school.csv',index=False) print(df)
目标爬取URL
https://aldrines.fcps.edu/staff-directory?keywords=&field_last_name_from=&field_last_name_to=&items_per_page=10&page= https://aldrines.fcps.edu
问题排查与修复
核心问题点
- 硬编码域名未全局替换:原代码中构造详情页URL时,硬写了
https://fairfaxhs.fcps.edu前缀,切换域名时仅修改了列表页URL,导致详情页请求的还是旧域名地址,自然无法获取数据。 - 异常捕获过于宽泛:
except: pass会吞掉所有错误(比如请求失败、元素找不到等),无法定位具体问题,调试难度大。
修改后的代码
import pandas as pd from bs4 import BeautifulSoup import requests # 只需修改此处域名前缀即可切换目标学校 school_domain = "aldrines.fcps.edu" base_url = f"https://{school_domain}" staff_list_url = f"{base_url}/staff-directory?field_last_name_from=&field_last_name_to=&items_per_page=10&keywords=&page={{page}}" data = [] for page in range(0, 2): try: # 请求列表页并检查HTTP状态 list_res = requests.get(staff_list_url.format(page=page)) list_res.raise_for_status() soup = BeautifulSoup(list_res.text, 'lxml') # 提取员工详情页链接 staff_cells = soup.table.select('tr td[class="views-field views-field-rendered-item"]') for cell in staff_cells: try: detail_href = cell.a.get('href') detail_url = f"{base_url}{detail_href}" # 请求详情页并检查状态 detail_res = requests.get(detail_url) detail_res.raise_for_status() soup2 = BeautifulSoup(detail_res.text, 'lxml') # 提取数据 name = soup2.select_one('h1.node__title.fcps-color--dark11').get_text(strip=True) position = soup2.select_one('h1+div').get_text(strip=True) data.append({ 'Name': name, 'Position': position, 'contact_url': detail_url }) except Exception as e: print(f"处理详情页失败 {detail_url}: {str(e)}") continue except Exception as e: print(f"请求列表页(页码{page})失败: {str(e)}") continue # 保存并输出结果 df = pd.DataFrame(data) df.to_csv('fcps_school.csv', index=False) print(df)
关键优化点
- 用
school_domain变量统一管理域名,切换目标学校时仅需修改一行代码 - 增加
raise_for_status()检查HTTP请求是否成功,避免请求失败后继续无效解析 - 拆分异常捕获逻辑,分别处理列表页和详情页的错误,并打印具体信息,便于调试
- 移除宽泛的
except: pass,精准定位问题
内容的提问来源于stack exchange,提问作者Syed
相关产品推荐
相关产品推荐

