SEC AAER站点批量下载关联文件时的403 Forbidden错误求助
SEC AAER站点批量下载关联文件时的403 Forbidden错误求助
我最近在尝试从SEC的AAER站点批量下载300个关联文件,其中大部分是PDF格式,但也有一些是需要另存为PDF的网页。我正在自学Python网络爬虫,本以为这个任务难度不大,但卡在了下载环节的403错误上,一直没法突破。
目前我已经成功实现了爬取文件链接和对应的4位编号(用来给下载后的文件命名),这部分代码运行正常:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time import os import requests # Set up Chrome options to allow direct PDF download (for the download step) download_path = "C:/Users/taylo/Downloads/sec_aaer_downloads" chrome_options = Options() chrome_options.add_experimental_option("prefs", { "download.default_directory": download_path, # Specify your preferred download directory "download.prompt_for_download": False, # Disable download prompt "plugins.always_open_pdf_externally": True, # Automatically open PDF in browser "safebrowsing.enabled": False, # Disable Chrome’s safe browsing check that can block downloads "profile.default_content_settings.popups": 0 # Disable popups }) # Set up the webdriver with options driver = webdriver.Chrome(executable_path="C:/chromedriver/chromedriver", options=chrome_options) # URLs for pages 1, 2, and 3 urls = [ "https://www.sec.gov/enforcement-litigation/accounting-auditing-enforcement-releases?page=0", "https://www.sec.gov/enforcement-litigation/accounting-auditing-enforcement-releases?page=1", "https://www.sec.gov/enforcement-litigation/accounting-auditing-enforcement-releases?page=2" ] # Initialize an empty list to store the URLs and AAER numbers pdf_data = [] # Loop through each URL (pages 1, 2, and 3) for url in urls: print(f"Scraping URL: {url}...") driver.get(url) # Wait for the table rows containing links to be loaded WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.XPATH, '//*[@id="block-uswds-sec-content"]/div/div/div[3]/div/table/tbody/tr[1]'))) # Extract the link and AAER number from each row on the current page rows = driver.find_elements(By.XPATH, '//*[@id="block-uswds-sec-content"]/div/div/div[3]/div/table/tbody/tr') for row in rows: try: # Extract the link from the first column (PDF link) link_element = row.find_element(By.XPATH, './/td[2]/div[1]/a') link_href = link_element.get_attribute('href') # Extract the AAER number from the second column aaer_text_element = row.find_element(By.XPATH, './/td[2]/div[2]/span[2]') aaer_text = aaer_text_element.text aaer_number = aaer_text.split("AAER-")[1].split()[0] # Extract the number after AAER- # Store the data in a list of dictionaries pdf_data.append({'link': link_href, 'aaer_number': aaer_number}) except Exception as e: print(f"Error extracting data from row: {e}") # Print the scraped data (optional for verification) for entry in pdf_data: print(f"Link: {entry['link']}, AAER Number: {entry['aaer_number']}")
但是当我尝试用以下代码进行下载时,始终无法突破403错误,文件根本下载不下来:
import os import time import requests # Set the download path download_path = "C:/Users/taylo/Downloads/sec_aaer_downloads" os.makedirs(download_path, exist_ok=True) # Loop through each entry in the pdf_data list for entry in pdf_data: try: # Extract the PDF link and AAER number link_href = entry['link'] aaer_number = entry['aaer_number'] # Send a GET request to download the PDF pdf_response = requests.get(link_href, stream=True, headers={ "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36" }) # Check if the request was successful if pdf_response.status_code == 200: # Save the PDF to the download folder, using the AAER number as the filename pdf_file_path = os.path.join(download_path, f"{aaer_number}.pdf") with open(pdf_file_path, "wb") as pdf_file: for chunk in pdf_response.iter_content(chunk_size=8192): pdf_file.write(chunk) print(f"Downloaded: {aaer_number}.pdf") else: print(f"Failed to download the file from {link_href}, status code: {pdf_response.status_code}") except Exception as e: print(f"Error downloading the PDF for AAER {aaer_number}: {e}")
这时候手动下载可能更快,但我更想搞清楚自己哪里出错了。我已经试过设置User-Agent请求头,也试过用Selenium模拟用户点击下载,但都没解决问题。希望能得到大家的建议!
备注:内容来源于stack exchange,提问作者Taylor James
相关产品推荐
相关产品推荐

