You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用wget下载PDF时遭遇HTTP Error 403: Forbidden问题求助

Hey, let's break down why you're hitting that 403 Forbidden error and fix it up. The core issue here is that even though you're using Selenium to navigate, your wget.download() calls are making independent requests that don't carry over the browser's session context—so the site's anti-bot systems are flagging them as suspicious.

The Root Cause

The Florida DEMS site's anti-bot measures are likely detecting that your wget requests don't match the context of a real browser session. When you use Selenium, the browser gets valid cookies and session tokens after loading the page, but wget sends fresh requests without those cookies. Plus, its default request headers might not fully mimic a browser (even if you set the User-Agent elsewhere). That's why standalone wget works for single files (maybe you're testing it right after browsing the site in your regular browser, which has cached cookies), but when run in the script, it's blocked.

Fixes to Try

Option 1: Reuse Selenium's Cookies in Download Requests

Grab the cookies from your Selenium driver and pass them to a more flexible library like requests (instead of wget) so your download requests look like they're part of the same session.

Here's the modified code:

import requests
import os
import wget
from selenium.webdriver.chrome.options import Options

chrome_options = Options()
chrome_options.add_argument("--incognito")
chrome_options.add_argument("--disable-plugins-discovery")
chrome_options.add_argument("--start-maximized")
driver = webdriver.Chrome(chrome_path, options=chrome_options)

# Create a requests session that inherits the browser's cookies
session = requests.Session()
for cookie in driver.get_cookies():
    session.cookies.set(cookie['name'], cookie['value'])

# Match your browser's headers exactly
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/85.0.4183.83 Safari/537.36',
    'Referer': driver.current_url  # Critical: tells the site where your request originates
}

print('Starting Data Download')
link_counter = 0
download_counter = 0
link_n = len(result_full) - 152
download_list = []

for links in result_full:
    if links.text.find('Data Report') > 0:
        link_url = links.get_attribute('href')
        filename = wget.filename_from_url(link_url)
        save_path = f'{pdf_output_path}/{filename}'
        
        if not os.path.exists(save_path):
            # Use the session to download with valid context
            response = session.get(link_url, headers=headers)
            response.raise_for_status()  # Catch HTTP errors early
            
            with open(save_path, 'wb') as f:
                f.write(response.content)
            
            download_counter += 1
            download_list.append(filename)  # Fixed: your original code missed the filename here
    
    print("Downloading", links.text)
    link_counter += 1
    print(f'{round((link_counter)*100/link_n,2)}% Complete')

print('Download of New Files Complete')
print(f'{download_counter} Files Created')

Option 2: Let Selenium Handle Downloads Directly

Configure Chrome to auto-download files to your target folder, so all requests stay within the browser context—this makes it way harder for anti-bot systems to detect your script.

import os
import wget
import time
from selenium.webdriver.chrome.options import Options

chrome_options = Options()
chrome_options.add_argument("--incognito")
chrome_options.add_argument("--disable-plugins-discovery")
chrome_options.add_argument("--start-maximized")

# Set up auto-download preferences
prefs = {
    "download.default_directory": pdf_output_path,
    "download.prompt_for_download": False,
    "download.directory_upgrade": True,
    "plugins.always_open_pdf_externally": True  # Skip PDF viewer, download directly
}
chrome_options.add_experimental_option("prefs", prefs)

driver = webdriver.Chrome(chrome_path, options=chrome_options)

print('Starting Data Download')
link_counter = 0
download_counter = 0
link_n = len(result_full) - 152
download_list = []

for links in result_full:
    if links.text.find('Data Report') > 0:
        link_url = links.get_attribute('href')
        filename = wget.filename_from_url(link_url)
        save_path = f'{pdf_output_path}/{filename}'
        
        if not os.path.exists(save_path):
            links.click()  # Trigger download via browser click
            time.sleep(2)  # Small delay to let the download initiate
            download_counter += 1
            download_list.append(filename)
    
    print("Downloading", links.text)
    link_counter += 1
    print(f'{round((link_counter)*100/link_n,2)}% Complete')

print('Download of New Files Complete')
print(f'{download_counter} Files Created')

Option 3: Add Rate Limiting

Even with valid session context, downloading too many files too fast can trigger flags. Add a small delay between requests to mimic human behavior:

import time
# Insert this after each download action in your loop
time.sleep(1)  # Adjust the duration based on how strict the site is

Quick Notes

  • I fixed the download_list.append line in your original code (it was missing the filename parameter).
  • Always make sure you're complying with the site's robots.txt and terms of service when scraping or downloading data.

内容的提问来源于stack exchange,提问作者Clovis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 23:22:29