网站批量下载PDF遇SSL证书验证失败问题求助
Hey there, let's break down why you're hitting that SSL error while trying to grab those COVID bulletin PDFs automatically. The core issue here is that the server hosting the PDFs (dl.bbmpgov.in) has an SSL certificate your system doesn't recognize as trusted—hence the certificate verify failed message. Your retry setup was a smart move, but there are two key gaps in your code that kept the error from going away:
- You created a
sessionwith retry logic, but didn't actually use it for most of your requests (including the critical PDF downloads). - You didn't tell
requestshow to handle the untrusted SSL certificate for the PDF host.
Let's fix this step by step.
Solution 1: Skip SSL Verification (Quick Personal Use Fix)
If this is just for your own personal use and you trust the source, you can temporarily disable SSL verification. Important: This isn't safe for production code, as it leaves you open to potential man-in-the-middle attacks.
Here's the modified code that uses your session (with retries) and skips SSL checks:
import os import requests from urllib.parse import urljoin from bs4 import BeautifulSoup from requests.adapters import HTTPAdapter from requests.packages.urllib3.util.retry import Retry url = "http://bbmp.gov.in/en/covid19bulletins" folder_location = r'C:\Users\maria.fh\Documents\Automatically downloaded files' # Create folder if it doesn't exist if not os.path.exists(folder_location): os.mkdir(folder_location) # Set up session with retry logic session = requests.Session() retry = Retry(connect=3, backoff_factor=0.8) adapter = HTTPAdapter(max_retries=retry) session.mount('http://', adapter) session.mount('https://', adapter) # Use the session to fetch the main page (instead of standalone requests.get) response = session.get(url) soup = BeautifulSoup(response.text, "html.parser") # Loop through PDF links and download using the session for link in soup.select("a[href$='.pdf']"): pdf_url = urljoin(url, link['href']) filename = os.path.join(folder_location, link['href'].split('/')[-1]) # Use session.get with verify=False to bypass SSL certificate check with open(filename, 'wb') as f: pdf_response = session.get(pdf_url, verify=False) f.write(pdf_response.content) print(f"Successfully downloaded: {filename}")
Key changes made:
- Replaced all standalone
requests.get()calls withsession.get()to leverage your pre-configured retry logic. - Added
verify=Falseto the PDF download request to bypass the certificate verification failure.
Solution 2: Trust the Server's Certificate (Recommended for Long-Term Use)
For a secure, long-term fix, you can download the SSL certificate for dl.bbmpgov.in and tell requests to explicitly trust it:
- Visit
https://dl.bbmpgov.inin your browser, download the certificate file (usually saved as.pemor.crt). - Update the
session.get()call for PDFs to point to your certificate file:pdf_response = session.get(pdf_url, verify="/path/to/your/certificate-file.pem")
This way, you're trusting only this specific server's certificate without disabling all SSL checks.
Why Your Original Code Didn't Work
- You created a session with retries, but used
requests.get(url)instead ofsession.get(url)for the main page. - When downloading PDFs, you used another standalone
requests.get()that didn't inherit the session's retry logic or SSL settings. - Retry logic only handles temporary network issues or timeouts—it can't fix SSL certificate verification failures.
内容的提问来源于stack exchange,提问作者maria francis

