Python批量爬取Gem招标网PDF时随机部分文件下载失败问题排查
问题根因
最终下载量仅为预期一半,核心是现有代码存在4个关键缺陷:
- 链接提取错误:列表页
.bid_no > a的href指向招标详情页,并非直接PDF资源,原代码直接将详情页的HTML响应写入PDF文件,必然产生大量无效文件 - 会话生命周期错误:收集链接的
with requests.Session()代码块结束后会话对象已被销毁,下载环节调用s.get属于非法上下文引用,部分环境下会直接触发请求失败 - 无反爬适配与容错机制:站点对无标识请求、高频请求会随机拦截返回错误响应,原代码无请求头伪装、无超时、无失败重试、无响应合法性校验,被拦截后直接写入错误内容
- 字典键覆盖风险:直接用招标编号作为字典key,若存在同编号重发招标会直接覆盖旧链接,导致漏爬
修正实现
核心调整点:
- 全程复用同一个Session,统一添加合法浏览器请求头
- 遍历列表页拿到详情页地址后,先请求详情页提取真实PDF下载链接(页面中class为
showbidDocument的a标签即为PDF入口,无需链接带.pdf后缀) - 增加3次失败重试、请求超时、0.8秒请求间隔,规避反爬拦截
- 增加文件存在判断,已下载的文件自动跳过,支持中断后续跑
- 增加响应类型校验,仅当响应为PDF格式时写入文件,避免存无效HTML内容
修正后可直接运行的代码:
import os import time import requests from bs4 import BeautifulSoup as bs # 配置项 end_page = 900 save_path = "./gem_bid_pdfs" # 替换为你的本地存储路径 os.makedirs(save_path, exist_ok=True) request_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", "Referer": "https://bidplus.gem.gov.in/bidlists" } max_retry = 3 request_interval = 0.8 # 单位秒,可根据实际响应情况调整 def request_with_retry(url, session): """带重试的请求封装""" for i in range(max_retry): try: resp = session.get(url, headers=request_headers, timeout=15) resp.raise_for_status() return resp except Exception as e: if i == max_retry -1: print(f"请求{url}失败,重试次数耗尽:{str(e)}") return None time.sleep(1) if __name__ == "__main__": current_page = 1 total_pdf_count = 0 with requests.Session() as s: # 先访问首页初始化cookie s.get("https://bidplus.gem.gov.in/bidlists", headers=request_headers, timeout=15) while current_page <= end_page: print(f"正在爬取第{current_page}页") list_url = f"https://bidplus.gem.gov.in/bidlists?bidlists&page_no={current_page}" list_resp = request_with_retry(list_url, s) if not list_resp: current_page +=1 continue list_soup = bs(list_resp.content, 'lxml') # 提取当前页所有详情页链接 bid_items = list_soup.select('.bid_no > a') for item in bid_items: bid_no = item.text.strip().replace('/', '_') save_file = os.path.join(save_path, f"{bid_no}.pdf") # 已下载的跳过 if os.path.exists(save_file) and os.path.getsize(save_file) > 1024: total_pdf_count +=1 continue detail_url = 'https://bidplus.gem.gov.in' + item['href'] detail_resp = request_with_retry(detail_url, s) if not detail_resp: time.sleep(request_interval) continue detail_soup = bs(detail_resp.content, 'lxml') # 提取真实PDF链接 pdf_tag = detail_soup.select_one('a.showbidDocument') if not pdf_tag: print(f"招标{bid_no}未找到PDF资源,可能已下架") time.sleep(request_interval) continue pdf_url = 'https://bidplus.gem.gov.in' + pdf_tag['href'] pdf_resp = request_with_retry(pdf_url, s) if not pdf_resp: time.sleep(request_interval) continue # 校验是否为PDF if 'application/pdf' not in pdf_resp.headers.get('Content-Type', ''): print(f"招标{bid_no}下载链接返回非PDF内容,跳过") time.sleep(request_interval) continue # 写入文件 with open(save_file, 'wb') as f: f.write(pdf_resp.content) total_pdf_count +=1 print(f"已成功下载第{total_pdf_count}份PDF:{bid_no}") time.sleep(request_interval) # 获取总页数 if current_page ==1: last_page_tag = list_soup.select_one('.pagination li:last-of-type > a') if last_page_tag: total_page = int(last_page_tag['data-ci-pagination-page']) end_page = min(end_page, total_page) print(f"站点总页数为{total_page}") current_page +=1 print(f"爬取完成,共成功下载{total_pdf_count}份PDF文件")
运行说明
- 首次运行前修改
save_path为你本地的实际存储路径即可 - 如果运行过程中出现大量请求失败,将
request_interval调整为1.5-2秒即可,无需额外配置代理 - 少数已下架的历史招标会自动跳过并打印日志,不会中断整体爬取任务
- 程序支持断点续跑,中途中断后重新运行会自动跳过已下载的文件,不需要从头开始
内容的提问来源于stack exchange,提问作者user17830341
相关产品推荐
相关产品推荐

