You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Requests和BeautifulSoup的Python PDF爬取优化问询

优化PDF检测爬虫:识别无.pdf后缀的PDF文件

你的代码核心思路没问题,但存在几个可以优化的点,尤其是URL拼接逻辑和Content-Type的检测方式,以下是调整后的完整代码,解决无后缀PDF的识别问题:

优化后的代码

import requests
from bs4 import BeautifulSoup
import pandas as pd
from urllib.parse import urljoin

# 目标页面URL
base_url = "https://machado.mec.gov.br/obra-completa-lista/itemlist/category/24-conto"

# 创建会话,提升请求效率并保持连接
session = requests.Session()
session.headers.update({
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
})

results = []
try:
    response = session.get(base_url, timeout=10)
    response.raise_for_status()  # 主动抛出HTTP错误

    soup = BeautifulSoup(response.content, "html.parser")
    links = soup.find_all("a")

    for link in links:
        href = link.get("href")
        if not href:
            continue
        
        # 处理相对/绝对路径,避免URL拼接错误
        full_url = urljoin(base_url, href)
        
        # 过滤非HTTP/HTTPS链接(如mailto、javascript)
        if not full_url.startswith(("http://", "https://")):
            continue

        is_pdf = False
        status_code = None
        content_type = None

        try:
            # 优先用HEAD请求获取头部信息,减少带宽消耗
            head_response = session.head(full_url, allow_redirects=True, timeout=10)
            status_code = head_response.status_code
            content_type = head_response.headers.get("Content-Type", "")
            
            # 宽松检测Content-Type,兼容带额外参数的情况(如application/pdf; charset=utf-8)
            if "application/pdf" in content_type.lower():
                is_pdf = True
            # 同时保留后缀检测作为补充
            elif href.lower().endswith(".pdf"):
                is_pdf = True

        except requests.exceptions.RequestException:
            # HEAD请求失败时, fallback到GET请求只获取头部
            try:
                get_response = session.get(full_url, stream=True, timeout=10)
                status_code = get_response.status_code
                content_type = get_response.headers.get("Content-Type", "")
                if "application/pdf" in content_type.lower():
                    is_pdf = True
                elif href.lower().endswith(".pdf"):
                    is_pdf = True
                # 关闭连接,避免资源泄漏
                get_response.close()
            except requests.exceptions.RequestException as e:
                status_code = "Request Failed"
                print(f"请求链接失败: {full_url}, 错误: {str(e)}")

        results.append({
            "Link": full_url,
            "Status": status_code,
            "Arquivo": href,
            "PDF": is_pdf
        })

    # 生成DataFrame
    df = pd.DataFrame(results)
    print(df)

except requests.exceptions.RequestException as e:
    print(f"请求基础页面失败: {str(e)}")

关键优化点说明

  • URL拼接修复:用urljoin替代直接字符串拼接,自动处理相对路径和绝对路径,避免出现类似https://xxx.comhttps://xxx.com/xxx的错误URL
  • Content-Type宽松检测:从严格等于application/pdf改为判断是否包含该字符串,兼容部分服务器返回的application/pdf; charset=utf-8这类带额外参数的响应头
  • 请求容错处理:HEAD请求可能被部分服务器拦截,增加GET请求的fallback逻辑,确保能获取到正确的Content-Type
  • 无效链接过滤:跳过mailto:、javascript:等非网页链接,减少无效请求
  • 会话复用:使用requests.Session()复用TCP连接,提升请求效率,同时添加User-Agent模拟浏览器,避免被服务器拦截
  • 错误捕获:添加全局和局部的异常捕获,避免单个链接请求失败导致整个脚本终止

内容的提问来源于stack exchange,提问作者Luiz Mário Andrade

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 13:13:07