实践《Python网络爬虫》代码时遇SSL证书验证失败错误求助
Hey there, I totally get how frustrating that SSL error can be when you're just trying to follow along with your web scraping book. Let's break down what's happening and fix it—here are a few solid solutions:
1. Quick Fix: Bypass SSL Verification (For Testing Only!)
The issue here is that Wikipedia automatically redirects your HTTP request to HTTPS, and your Python environment is failing to verify the SSL certificate. You can temporarily skip this check with urllib, but never use this in production code (it's insecure):
from urllib.request import urlopen from urllib.error import URLError import ssl from bs4 import BeautifulSoup import re pages = set() # Create an SSL context that skips certificate verification ctx = ssl.create_default_context() ctx.check_hostname = False ctx.verify_mode = ssl.CERT_NONE def getLinks(pageUrl): global pages try: # Pass the custom context to urlopen html = urlopen("http://en.wikipedia.org" + pageUrl, context=ctx) bsObj = BeautifulSoup(html, "html.parser") for link in bsObj.findAll("a", href=re.compile("^(/wiki/)")): if 'href' in link.attrs and link.attrs['href'] not in pages: newPage = link.attrs['href'] print(newPage) pages.add(newPage) getLinks(newPage) except URLError as e: print(f"Failed to access page: {e}") # Start scraping from Wikipedia's homepage getLinks("")
2. Better Solution: Use the requests Library
requests handles SSL certificates more smoothly out of the box, and it makes your code cleaner. First install it if you haven't:
pip install requests
Then rewrite your code like this:
import requests from bs4 import BeautifulSoup import re pages = set() def getLinks(pageUrl): global pages try: # Requests automatically handles HTTPS redirects and certificate checks response = requests.get("http://en.wikipedia.org" + pageUrl) response.raise_for_status() # Raise an error if the request fails bsObj = BeautifulSoup(response.text, "html.parser") for link in bsObj.findAll("a", href=re.compile("^(/wiki/)")): if 'href' in link.attrs and link.attrs['href'] not in pages: newPage = link.attrs['href'] print(newPage) pages.add(newPage) getLinks(newPage) except requests.exceptions.RequestException as e: print(f"Failed to access page: {e}") getLinks("")
If you still hit certificate issues here, you can add verify=False to the get() call (again, only for testing!).
3. Most Secure Fix: Update Your Python SSL Certificates
The root cause might be that your Python installation is missing trusted root certificates. This is common on macOS. To fix it permanently:
- Find your Python installation folder (e.g.,
/Applications/Python 3.10/) - Run the
Install Certificates.commandscript inside that folder - Or run this command in your terminal:
/Applications/Python\ 3.x/Install\ Certificates.command
Replace 3.x with your actual Python version. This will install the necessary root certificates so all SSL requests work correctly moving forward.
A Quick Note
Remember to respect Wikipedia's robots.txt and don't scrape too aggressively—add delays between requests if you're crawling a lot of pages to avoid getting your IP blocked!
内容的提问来源于stack exchange,提问作者Catherine4j

