You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫报错TypeError: __init__() got an unexpected keyword argument 'proxies'——代理参数配置位置咨询

Fixing the TypeError: __init__() got an unexpected keyword argument 'proxies' in Your Web Scraper

Let's break down why you're hitting this error first: urllib.request.Request doesn't accept a proxies parameter—that's a feature exclusive to the requests library, not urllib. You're mixing two different HTTP client libraries here, which is causing the conflict.

Here are two clean solutions to fix this, depending on whether you want to stick with urllib or switch fully to requests (which I recommend for its simpler, more intuitive API):

You already imported requests and set up a Session, so let's lean into that. Sessions make it trivial to persist headers and proxies across requests. Here's how to rewrite your code properly:

# Import only what you need for cleaner code
from bs4 import BeautifulSoup
import requests

# Proxy configuration note: Your original dict had duplicate "https" keys—Python only keeps the last one!
# Pick one valid proxy, or set up rotation logic if you need multiple
PROXY = {"https": "88.198.26.145:8080"}
url = "https://www.ebay.de/b/Laptops-Notebooks/175672/bn_1618754?LH_ItemCondition=7000&mag=1&rt=nc&_dmd=1&_sop=1"

# Use a requests Session to manage proxies and headers consistently
with requests.Session() as session:
    # Apply headers and proxies to the session
    session.headers.update({"User-Agent": "Mozilla/5.0"})
    session.proxies.update(PROXY)
    
    try:
        # Fetch the webpage
        response = session.get(url)
        response.raise_for_status()  # Catch HTTP errors like 404/500
        
        # Parse the HTML with BeautifulSoup
        soup = BeautifulSoup(response.text, "html5lib")
        
        # Your data collection logic here (make sure to define 'shipping' first!)
        # shipping = soup.find("some-selector")
        # print(shipping.text.strip())
        
        all_data = []  # Initialize your data storage inside the session context
        
    except requests.exceptions.RequestException as e:
        print(f"Request failed: {e}")

Option 2: Stick with urllib

If you prefer to keep using urllib, you need to set up a proxy handler explicitly instead of passing a proxies parameter to Request. Here's how to adjust your code:

from urllib.request import ProxyHandler, Request, build_opener, install_opener
from bs4 import BeautifulSoup

PROXY = {"https": "88.198.26.145:8080"}
url = "https://www.ebay.de/b/Laptops-Notebooks/175672/bn_1618754?LH_ItemCondition=7000&mag=1&rt=nc&_dmd=1&_sop=1"

# Create a proxy handler and build a custom opener
proxy_handler = ProxyHandler(PROXY)
opener = build_opener(proxy_handler)
install_opener(opener)  # Make this opener the default for urlopen

# Now send the request with headers (no proxies parameter needed!)
req = Request(url, headers={"User-Agent": "Mozilla/5.0"})
webpage = urlopen(req).read()

# Parse and collect data
soup = BeautifulSoup(webpage, "html5lib")
# shipping = soup.find("some-selector")
# print(shipping.text.strip())

Key Takeaway

The core mistake was mixing urllib's Request class with requests-style proxies syntax. Pick one library and stick to its API—requests is almost always the better choice for web scraping due to its simpler, more feature-rich interface.

内容的提问来源于stack exchange,提问作者Tretecou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 06:22:34