Python请求Longman在线词典网站时出现requests.exceptions.ConnectionError的解决方法咨询
requests.exceptions.ConnectionError When Scraping Longman Online Dictionary Hey, I’ve run into this exact issue with dictionary sites before—they’re pretty strict about blocking non-browser requests. Since your code works with other URLs, the problem is definitely Longman’s anti-scraping measures flagging your request as suspicious. Let’s walk through the fixes step by step:
1. Add a Proper User-Agent Header (Critical!)
Longman’s server checks the User-Agent string to identify if the request is coming from a browser or a script. The default requests User-Agent is easy to spot. Replace it with a real browser’s User-Agent to mimic a normal visit:
from requests import get from bs4 import BeautifulSoup URL = "https://www.ldoceonline.com/dictionary/" # Mimic a Chrome browser on Windows headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } try: response = get(URL, headers=headers) response.raise_for_status() # Raise an error if the request fails (e.g., 403, 404) print("Request successful!") # Now you can parse with BeautifulSoup: soup = BeautifulSoup(response.text, "html.parser") except Exception as e: print(f"Error: {e}")
2. Add a Timeout to Avoid Hanging Connections
Sometimes the server might stall instead of responding, leading to a connection abort. Adding a timeout ensures your request doesn’t wait indefinitely:
response = get(URL, headers=headers, timeout=10) # Wait max 10 seconds for a response
3. Slow Down Your Requests
If you’re making multiple requests in quick succession, Longman will block you immediately. Add a delay between requests using time.sleep():
import time # After each request: time.sleep(2) # Wait 2 seconds before the next request
4. Use a Session for Persistent Connections
Using requests.Session() maintains cookies and connection settings across requests, which can make your traffic look more like a real user’s:
from requests import Session session = Session() session.headers.update(headers) # Apply headers to all requests in the session response = session.get(URL, timeout=10)
Why You Got That Error
The RemoteDisconnected('Remote end closed connection without response') message means Longman’s server detected your request as non-human and cut the connection without sending any data. Adding the User-Agent header should resolve this for most cases—this is the most common fix for this type of error with dictionary sites.
内容的提问来源于stack exchange,提问作者AmirhosseinHG

