You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

实践《Python网络爬虫》代码时遇SSL证书验证失败错误求助

Fixing SSL: CERTIFICATE_VERIFY_FAILED When Scraping Wikipedia

Hey there, I totally get how frustrating that SSL error can be when you're just trying to follow along with your web scraping book. Let's break down what's happening and fix it—here are a few solid solutions:

1. Quick Fix: Bypass SSL Verification (For Testing Only!)

The issue here is that Wikipedia automatically redirects your HTTP request to HTTPS, and your Python environment is failing to verify the SSL certificate. You can temporarily skip this check with urllib, but never use this in production code (it's insecure):

from urllib.request import urlopen
from urllib.error import URLError
import ssl
from bs4 import BeautifulSoup
import re

pages = set()

# Create an SSL context that skips certificate verification
ctx = ssl.create_default_context()
ctx.check_hostname = False
ctx.verify_mode = ssl.CERT_NONE

def getLinks(pageUrl):
    global pages
    try:
        # Pass the custom context to urlopen
        html = urlopen("http://en.wikipedia.org" + pageUrl, context=ctx)
        bsObj = BeautifulSoup(html, "html.parser")
        
        for link in bsObj.findAll("a", href=re.compile("^(/wiki/)")):
            if 'href' in link.attrs and link.attrs['href'] not in pages:
                newPage = link.attrs['href']
                print(newPage)
                pages.add(newPage)
                getLinks(newPage)
    except URLError as e:
        print(f"Failed to access page: {e}")

# Start scraping from Wikipedia's homepage
getLinks("")

2. Better Solution: Use the requests Library

requests handles SSL certificates more smoothly out of the box, and it makes your code cleaner. First install it if you haven't:

pip install requests

Then rewrite your code like this:

import requests
from bs4 import BeautifulSoup
import re

pages = set()

def getLinks(pageUrl):
    global pages
    try:
        # Requests automatically handles HTTPS redirects and certificate checks
        response = requests.get("http://en.wikipedia.org" + pageUrl)
        response.raise_for_status() # Raise an error if the request fails
        
        bsObj = BeautifulSoup(response.text, "html.parser")
        for link in bsObj.findAll("a", href=re.compile("^(/wiki/)")):
            if 'href' in link.attrs and link.attrs['href'] not in pages:
                newPage = link.attrs['href']
                print(newPage)
                pages.add(newPage)
                getLinks(newPage)
    except requests.exceptions.RequestException as e:
        print(f"Failed to access page: {e}")

getLinks("")

If you still hit certificate issues here, you can add verify=False to the get() call (again, only for testing!).

3. Most Secure Fix: Update Your Python SSL Certificates

The root cause might be that your Python installation is missing trusted root certificates. This is common on macOS. To fix it permanently:

  • Find your Python installation folder (e.g., /Applications/Python 3.10/)
  • Run the Install Certificates.command script inside that folder
  • Or run this command in your terminal:
/Applications/Python\ 3.x/Install\ Certificates.command

Replace 3.x with your actual Python version. This will install the necessary root certificates so all SSL requests work correctly moving forward.

A Quick Note

Remember to respect Wikipedia's robots.txt and don't scrape too aggressively—add delays between requests if you're crawling a lot of pages to avoid getting your IP blocked!

内容的提问来源于stack exchange,提问作者Catherine4j

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:16:03