You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas和BeautifulSoup从指定HTTPS网站下载PDF失败求助

PDF下载脚本失效问题排查与解决

问题描述

尝试使用BeautifulSoup从指定网站下载PDF文件,现有脚本在示例网站正常运行,但在目标网站(REO家庭医生页面)无任何文件下载。目标页面地址:https://www.gems.gov.za/Healthcare-Providers/GEMS-Netwrk-of-Healthcare-Providers/Primary-Network/Family-Practitioners/REO-Family-Practitioners

使用的原脚本:

# Import libraries
import requests
from bs4 import BeautifulSoup

# URL from which pdfs to be downloaded
url = "https://www.gems.gov.za/Healthcare-Providers/GEMS-Netwrk-of-Healthcare-Providers/Specialist-Network/Obstetricians-and-gynaecologists-list/"

# Requests URL and get response object
response = requests.get(url)

# Parse text obtained
soup = BeautifulSoup(response.text, 'html.parser')

# Find all hyperlinks present on webpage
links = soup.find_all('a')

i = 0

# From all links check for pdf link and
# if present download file
for link in links:
    if ('.pdf' in link.get('href', [])):
        i += 1
        print("Downloading file: ", i)

        # Get response object for link
        response = requests.get(link.get('href'))

        # Write content in pdf file
        pdf = open("pdf"+str(i)+".pdf", 'wb')
        pdf.write(response.content)
        pdf.close()
        print("File ", i, " downloaded")

print("All PDF files downloaded")

核心问题分析

  1. URL错误:原脚本中填写的是妇产科医生列表页面,并非目标家庭医生页面,导致爬取对象错误。
  2. 动态内容加载:目标页面的PDF列表可能通过JavaScript动态渲染,requests仅能获取静态HTML,无法捕获动态生成的链接。
  3. 反爬拦截:无请求头的requests请求易被网站识别为爬虫,返回空页面或无效内容。
  4. 路径处理缺失:若PDF链接为相对路径,直接请求会导致404错误,需拼接完整URL。

修复后的脚本

方案1:静态页面适配(带请求头+路径处理)

适用于目标页面PDF链接为静态渲染的情况:

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

# 目标页面URL
base_url = "https://www.gems.gov.za/Healthcare-Providers/GEMS-Netwrk-of-Healthcare-Providers/Primary-Network/Family-Practitioners/REO-Family-Practitioners"

# 模拟浏览器请求头,避免反爬
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

# 获取页面内容,确保编码正确
response = requests.get(base_url, headers=headers)
response.encoding = response.apparent_encoding

# 解析页面
soup = BeautifulSoup(response.text, 'html.parser')
links = soup.find_all('a', href=True)

file_count = 0
for link in links:
    href = link['href']
    # 匹配PDF链接(忽略大小写)
    if href.lower().endswith('.pdf'):
        file_count += 1
        # 拼接完整URL,处理相对路径
        full_pdf_url = urljoin(base_url, href)
        print(f"正在下载文件: {file_count}")
        
        # 下载PDF文件
        pdf_response = requests.get(full_pdf_url, headers=headers)
        with open(f"pdf_{file_count}.pdf", 'wb') as f:
            f.write(pdf_response.content)
        print(f"文件 {file_count} 下载完成")

print("所有PDF文件下载完成")

方案2:动态页面适配(使用Selenium)

若目标页面PDF链接为动态生成,需用Selenium模拟浏览器加载完整页面:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
from urllib.parse import urljoin
import requests

# 配置Chrome无头模式(无界面运行)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("--disable-gpu")
chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")

# 初始化浏览器
driver = webdriver.Chrome(options=chrome_options)
base_url = "https://www.gems.gov.za/Healthcare-Providers/GEMS-Netwrk-of-Healthcare-Providers/Primary-Network/Family-Practitioners/REO-Family-Practitioners"

# 加载页面并等待动态内容渲染
driver.get(base_url)
driver.implicitly_wait(10)  # 等待10秒确保内容加载完成

# 获取完整页面源码
page_source = driver.page_source
driver.quit()

# 解析页面并下载PDF
soup = BeautifulSoup(page_source, 'html.parser')
links = soup.find_all('a', href=True)

file_count = 0
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

for link in links:
    href = link['href']
    if href.lower().endswith('.pdf'):
        file_count += 1
        full_pdf_url = urljoin(base_url, href)
        print(f"正在下载文件: {file_count}")
        pdf_response = requests.get(full_pdf_url, headers=headers)
        with open(f"pdf_{file_count}.pdf", 'wb') as f:
            f.write(pdf_response.content)
        print(f"文件 {file_count} 下载完成")

print("所有PDF文件下载完成")

内容的提问来源于stack exchange,提问作者Rodemire Tarazone

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 15:32:20