Python爬虫报错TypeError: __init__() got an unexpected keyword argument 'proxies'——代理参数配置位置咨询
TypeError: __init__() got an unexpected keyword argument 'proxies' in Your Web Scraper Let's break down why you're hitting this error first: urllib.request.Request doesn't accept a proxies parameter—that's a feature exclusive to the requests library, not urllib. You're mixing two different HTTP client libraries here, which is causing the conflict.
Here are two clean solutions to fix this, depending on whether you want to stick with urllib or switch fully to requests (which I recommend for its simpler, more intuitive API):
Option 1: Use requests (Cleaner, Recommended)
You already imported requests and set up a Session, so let's lean into that. Sessions make it trivial to persist headers and proxies across requests. Here's how to rewrite your code properly:
# Import only what you need for cleaner code from bs4 import BeautifulSoup import requests # Proxy configuration note: Your original dict had duplicate "https" keys—Python only keeps the last one! # Pick one valid proxy, or set up rotation logic if you need multiple PROXY = {"https": "88.198.26.145:8080"} url = "https://www.ebay.de/b/Laptops-Notebooks/175672/bn_1618754?LH_ItemCondition=7000&mag=1&rt=nc&_dmd=1&_sop=1" # Use a requests Session to manage proxies and headers consistently with requests.Session() as session: # Apply headers and proxies to the session session.headers.update({"User-Agent": "Mozilla/5.0"}) session.proxies.update(PROXY) try: # Fetch the webpage response = session.get(url) response.raise_for_status() # Catch HTTP errors like 404/500 # Parse the HTML with BeautifulSoup soup = BeautifulSoup(response.text, "html5lib") # Your data collection logic here (make sure to define 'shipping' first!) # shipping = soup.find("some-selector") # print(shipping.text.strip()) all_data = [] # Initialize your data storage inside the session context except requests.exceptions.RequestException as e: print(f"Request failed: {e}")
Option 2: Stick with urllib
If you prefer to keep using urllib, you need to set up a proxy handler explicitly instead of passing a proxies parameter to Request. Here's how to adjust your code:
from urllib.request import ProxyHandler, Request, build_opener, install_opener from bs4 import BeautifulSoup PROXY = {"https": "88.198.26.145:8080"} url = "https://www.ebay.de/b/Laptops-Notebooks/175672/bn_1618754?LH_ItemCondition=7000&mag=1&rt=nc&_dmd=1&_sop=1" # Create a proxy handler and build a custom opener proxy_handler = ProxyHandler(PROXY) opener = build_opener(proxy_handler) install_opener(opener) # Make this opener the default for urlopen # Now send the request with headers (no proxies parameter needed!) req = Request(url, headers={"User-Agent": "Mozilla/5.0"}) webpage = urlopen(req).read() # Parse and collect data soup = BeautifulSoup(webpage, "html5lib") # shipping = soup.find("some-selector") # print(shipping.text.strip())
Key Takeaway
The core mistake was mixing urllib's Request class with requests-style proxies syntax. Pick one library and stick to its API—requests is almost always the better choice for web scraping due to its simpler, more feature-rich interface.
内容的提问来源于stack exchange,提问作者Tretecou

