You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python批量爬取Gem招标网PDF时随机部分文件下载失败问题排查

问题根因

最终下载量仅为预期一半,核心是现有代码存在4个关键缺陷:

  • 链接提取错误:列表页.bid_no > a的href指向招标详情页,并非直接PDF资源,原代码直接将详情页的HTML响应写入PDF文件,必然产生大量无效文件
  • 会话生命周期错误:收集链接的with requests.Session()代码块结束后会话对象已被销毁,下载环节调用s.get属于非法上下文引用,部分环境下会直接触发请求失败
  • 无反爬适配与容错机制:站点对无标识请求、高频请求会随机拦截返回错误响应,原代码无请求头伪装、无超时、无失败重试、无响应合法性校验,被拦截后直接写入错误内容
  • 字典键覆盖风险:直接用招标编号作为字典key,若存在同编号重发招标会直接覆盖旧链接,导致漏爬
修正实现

核心调整点:

  • 全程复用同一个Session,统一添加合法浏览器请求头
  • 遍历列表页拿到详情页地址后,先请求详情页提取真实PDF下载链接(页面中class为showbidDocument的a标签即为PDF入口,无需链接带.pdf后缀)
  • 增加3次失败重试、请求超时、0.8秒请求间隔,规避反爬拦截
  • 增加文件存在判断,已下载的文件自动跳过,支持中断后续跑
  • 增加响应类型校验,仅当响应为PDF格式时写入文件,避免存无效HTML内容

修正后可直接运行的代码:

import os
import time
import requests
from bs4 import BeautifulSoup as bs

# 配置项
end_page = 900
save_path = "./gem_bid_pdfs"  # 替换为你的本地存储路径
os.makedirs(save_path, exist_ok=True)
request_headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
    "Referer": "https://bidplus.gem.gov.in/bidlists"
}
max_retry = 3
request_interval = 0.8  # 单位秒,可根据实际响应情况调整

def request_with_retry(url, session):
    """带重试的请求封装"""
    for i in range(max_retry):
        try:
            resp = session.get(url, headers=request_headers, timeout=15)
            resp.raise_for_status()
            return resp
        except Exception as e:
            if i == max_retry -1:
                print(f"请求{url}失败,重试次数耗尽:{str(e)}")
                return None
            time.sleep(1)

if __name__ == "__main__":
    current_page = 1
    total_pdf_count = 0
    with requests.Session() as s:
        # 先访问首页初始化cookie
        s.get("https://bidplus.gem.gov.in/bidlists", headers=request_headers, timeout=15)
        while current_page <= end_page:
            print(f"正在爬取第{current_page}页")
            list_url = f"https://bidplus.gem.gov.in/bidlists?bidlists&page_no={current_page}"
            list_resp = request_with_retry(list_url, s)
            if not list_resp:
                current_page +=1
                continue
            list_soup = bs(list_resp.content, 'lxml')
            
            # 提取当前页所有详情页链接
            bid_items = list_soup.select('.bid_no > a')
            for item in bid_items:
                bid_no = item.text.strip().replace('/', '_')
                save_file = os.path.join(save_path, f"{bid_no}.pdf")
                # 已下载的跳过
                if os.path.exists(save_file) and os.path.getsize(save_file) > 1024:
                    total_pdf_count +=1
                    continue
                detail_url = 'https://bidplus.gem.gov.in' + item['href']
                detail_resp = request_with_retry(detail_url, s)
                if not detail_resp:
                    time.sleep(request_interval)
                    continue
                detail_soup = bs(detail_resp.content, 'lxml')
                # 提取真实PDF链接
                pdf_tag = detail_soup.select_one('a.showbidDocument')
                if not pdf_tag:
                    print(f"招标{bid_no}未找到PDF资源,可能已下架")
                    time.sleep(request_interval)
                    continue
                pdf_url = 'https://bidplus.gem.gov.in' + pdf_tag['href']
                pdf_resp = request_with_retry(pdf_url, s)
                if not pdf_resp:
                    time.sleep(request_interval)
                    continue
                # 校验是否为PDF
                if 'application/pdf' not in pdf_resp.headers.get('Content-Type', ''):
                    print(f"招标{bid_no}下载链接返回非PDF内容,跳过")
                    time.sleep(request_interval)
                    continue
                # 写入文件
                with open(save_file, 'wb') as f:
                    f.write(pdf_resp.content)
                total_pdf_count +=1
                print(f"已成功下载第{total_pdf_count}份PDF:{bid_no}")
                time.sleep(request_interval)
            
            # 获取总页数
            if current_page ==1:
                last_page_tag = list_soup.select_one('.pagination li:last-of-type > a')
                if last_page_tag:
                    total_page = int(last_page_tag['data-ci-pagination-page'])
                    end_page = min(end_page, total_page)
                    print(f"站点总页数为{total_page}")
            current_page +=1
    print(f"爬取完成,共成功下载{total_pdf_count}份PDF文件")
运行说明
  • 首次运行前修改save_path为你本地的实际存储路径即可
  • 如果运行过程中出现大量请求失败,将request_interval调整为1.5-2秒即可,无需额外配置代理
  • 少数已下架的历史招标会自动跳过并打印日志,不会中断整体爬取任务
  • 程序支持断点续跑,中途中断后重新运行会自动跳过已下载的文件,不需要从头开始

内容的提问来源于stack exchange,提问作者user17830341

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.20 16:15:49