基于Requests和BeautifulSoup的Python PDF爬取优化问询
优化PDF检测爬虫:识别无
.pdf后缀的PDF文件 你的代码核心思路没问题,但存在几个可以优化的点,尤其是URL拼接逻辑和Content-Type的检测方式,以下是调整后的完整代码,解决无后缀PDF的识别问题:
优化后的代码
import requests from bs4 import BeautifulSoup import pandas as pd from urllib.parse import urljoin # 目标页面URL base_url = "https://machado.mec.gov.br/obra-completa-lista/itemlist/category/24-conto" # 创建会话,提升请求效率并保持连接 session = requests.Session() session.headers.update({ "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" }) results = [] try: response = session.get(base_url, timeout=10) response.raise_for_status() # 主动抛出HTTP错误 soup = BeautifulSoup(response.content, "html.parser") links = soup.find_all("a") for link in links: href = link.get("href") if not href: continue # 处理相对/绝对路径,避免URL拼接错误 full_url = urljoin(base_url, href) # 过滤非HTTP/HTTPS链接(如mailto、javascript) if not full_url.startswith(("http://", "https://")): continue is_pdf = False status_code = None content_type = None try: # 优先用HEAD请求获取头部信息,减少带宽消耗 head_response = session.head(full_url, allow_redirects=True, timeout=10) status_code = head_response.status_code content_type = head_response.headers.get("Content-Type", "") # 宽松检测Content-Type,兼容带额外参数的情况(如application/pdf; charset=utf-8) if "application/pdf" in content_type.lower(): is_pdf = True # 同时保留后缀检测作为补充 elif href.lower().endswith(".pdf"): is_pdf = True except requests.exceptions.RequestException: # HEAD请求失败时, fallback到GET请求只获取头部 try: get_response = session.get(full_url, stream=True, timeout=10) status_code = get_response.status_code content_type = get_response.headers.get("Content-Type", "") if "application/pdf" in content_type.lower(): is_pdf = True elif href.lower().endswith(".pdf"): is_pdf = True # 关闭连接,避免资源泄漏 get_response.close() except requests.exceptions.RequestException as e: status_code = "Request Failed" print(f"请求链接失败: {full_url}, 错误: {str(e)}") results.append({ "Link": full_url, "Status": status_code, "Arquivo": href, "PDF": is_pdf }) # 生成DataFrame df = pd.DataFrame(results) print(df) except requests.exceptions.RequestException as e: print(f"请求基础页面失败: {str(e)}")
关键优化点说明
- URL拼接修复:用
urljoin替代直接字符串拼接,自动处理相对路径和绝对路径,避免出现类似https://xxx.comhttps://xxx.com/xxx的错误URL - Content-Type宽松检测:从严格等于
application/pdf改为判断是否包含该字符串,兼容部分服务器返回的application/pdf; charset=utf-8这类带额外参数的响应头 - 请求容错处理:HEAD请求可能被部分服务器拦截,增加GET请求的fallback逻辑,确保能获取到正确的Content-Type
- 无效链接过滤:跳过
mailto:、javascript:等非网页链接,减少无效请求 - 会话复用:使用
requests.Session()复用TCP连接,提升请求效率,同时添加User-Agent模拟浏览器,避免被服务器拦截 - 错误捕获:添加全局和局部的异常捕获,避免单个链接请求失败导致整个脚本终止
内容的提问来源于stack exchange,提问作者Luiz Mário Andrade
相关产品推荐
相关产品推荐

