使用Pandas和BeautifulSoup从指定HTTPS网站下载PDF失败求助
PDF下载脚本失效问题排查与解决
问题描述
尝试使用BeautifulSoup从指定网站下载PDF文件,现有脚本在示例网站正常运行,但在目标网站(REO家庭医生页面)无任何文件下载。目标页面地址:https://www.gems.gov.za/Healthcare-Providers/GEMS-Netwrk-of-Healthcare-Providers/Primary-Network/Family-Practitioners/REO-Family-Practitioners
使用的原脚本:
# Import libraries import requests from bs4 import BeautifulSoup # URL from which pdfs to be downloaded url = "https://www.gems.gov.za/Healthcare-Providers/GEMS-Netwrk-of-Healthcare-Providers/Specialist-Network/Obstetricians-and-gynaecologists-list/" # Requests URL and get response object response = requests.get(url) # Parse text obtained soup = BeautifulSoup(response.text, 'html.parser') # Find all hyperlinks present on webpage links = soup.find_all('a') i = 0 # From all links check for pdf link and # if present download file for link in links: if ('.pdf' in link.get('href', [])): i += 1 print("Downloading file: ", i) # Get response object for link response = requests.get(link.get('href')) # Write content in pdf file pdf = open("pdf"+str(i)+".pdf", 'wb') pdf.write(response.content) pdf.close() print("File ", i, " downloaded") print("All PDF files downloaded")
核心问题分析
- URL错误:原脚本中填写的是妇产科医生列表页面,并非目标家庭医生页面,导致爬取对象错误。
- 动态内容加载:目标页面的PDF列表可能通过JavaScript动态渲染,
requests仅能获取静态HTML,无法捕获动态生成的链接。 - 反爬拦截:无请求头的
requests请求易被网站识别为爬虫,返回空页面或无效内容。 - 路径处理缺失:若PDF链接为相对路径,直接请求会导致404错误,需拼接完整URL。
修复后的脚本
方案1:静态页面适配(带请求头+路径处理)
适用于目标页面PDF链接为静态渲染的情况:
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin # 目标页面URL base_url = "https://www.gems.gov.za/Healthcare-Providers/GEMS-Netwrk-of-Healthcare-Providers/Primary-Network/Family-Practitioners/REO-Family-Practitioners" # 模拟浏览器请求头,避免反爬 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # 获取页面内容,确保编码正确 response = requests.get(base_url, headers=headers) response.encoding = response.apparent_encoding # 解析页面 soup = BeautifulSoup(response.text, 'html.parser') links = soup.find_all('a', href=True) file_count = 0 for link in links: href = link['href'] # 匹配PDF链接(忽略大小写) if href.lower().endswith('.pdf'): file_count += 1 # 拼接完整URL,处理相对路径 full_pdf_url = urljoin(base_url, href) print(f"正在下载文件: {file_count}") # 下载PDF文件 pdf_response = requests.get(full_pdf_url, headers=headers) with open(f"pdf_{file_count}.pdf", 'wb') as f: f.write(pdf_response.content) print(f"文件 {file_count} 下载完成") print("所有PDF文件下载完成")
方案2:动态页面适配(使用Selenium)
若目标页面PDF链接为动态生成,需用Selenium模拟浏览器加载完整页面:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup from urllib.parse import urljoin import requests # 配置Chrome无头模式(无界面运行) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") # 初始化浏览器 driver = webdriver.Chrome(options=chrome_options) base_url = "https://www.gems.gov.za/Healthcare-Providers/GEMS-Netwrk-of-Healthcare-Providers/Primary-Network/Family-Practitioners/REO-Family-Practitioners" # 加载页面并等待动态内容渲染 driver.get(base_url) driver.implicitly_wait(10) # 等待10秒确保内容加载完成 # 获取完整页面源码 page_source = driver.page_source driver.quit() # 解析页面并下载PDF soup = BeautifulSoup(page_source, 'html.parser') links = soup.find_all('a', href=True) file_count = 0 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } for link in links: href = link['href'] if href.lower().endswith('.pdf'): file_count += 1 full_pdf_url = urljoin(base_url, href) print(f"正在下载文件: {file_count}") pdf_response = requests.get(full_pdf_url, headers=headers) with open(f"pdf_{file_count}.pdf", 'wb') as f: f.write(pdf_response.content) print(f"文件 {file_count} 下载完成") print("所有PDF文件下载完成")
内容的提问来源于stack exchange,提问作者Rodemire Tarazone
相关产品推荐
相关产品推荐

