You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取网站遇连接错误,求解决方案

Troubleshooting Connection Errors When Web Scraping with Requests & BeautifulSoup

Hey there, let's work through this connection error you're facing while scraping with requests and BeautifulSoup. The "Remote end closed connection without response" error usually means the target website is rejecting your request—either because it detects you're a bot, or there's a network/configuration issue on your end. Here are some actionable fixes to try:

1. Fix Your Request Headers (Make Them Look Human)

Your initial headers were incomplete or unrealistic (like using 'Referer': 'hello' which is obviously not a valid referrer). Websites check these headers to verify you're a real browser. Let's build a more complete header set:

import requests
from bs4 import BeautifulSoup

url = 'https://www.example.com/bangalore/restaurants'
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://www.example.com/',  # Use the actual homepage of the target site
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
}

try:
    response = requests.get(url, headers=headers, timeout=10)
    response.raise_for_status()  # Catch HTTP errors like 403/404
    soup = BeautifulSoup(response.text, 'html.parser')
    print("Page fetched successfully!")
    # Add your parsing logic here
except requests.exceptions.ConnectionError as e:
    print(f"Connection failed: {e}")
except requests.exceptions.Timeout:
    print("Request timed out—try increasing the timeout value or checking your network")
except requests.exceptions.HTTPError as e:
    print(f"HTTP error occurred: {e}")

Key notes:

  • Use a realistic User-Agent (you can copy yours from your browser's dev tools)
  • Set the Referer to the actual homepage of the target site, not a random string
  • Add timeout to avoid hanging indefinitely if the site doesn't respond

2. Check If the Site Uses Dynamic Content (JS Rendering)

If the page loads most of its content via JavaScript (you can check by disabling JS in your browser and reloading the page), BeautifulSoup won't see that content—because requests only fetches the initial HTML. In this case, you'll need a tool that can execute JavaScript:

Using Selenium (Headless Browser)

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup

# Configure headless Chrome to avoid opening a browser window
options = Options()
options.add_argument('--headless=new')
options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
options.add_argument('--disable-blink-features=AutomationControlled')  # Bypass bot detection

driver = webdriver.Chrome(options=options)
try:
    driver.get('https://www.example.com/bangalore/restaurants')
    # Wait a few seconds for JS content to load (adjust as needed)
    driver.implicitly_wait(5)
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    print("Dynamic content fetched successfully!")
    # Parse the soup as usual
finally:
    driver.quit()  # Always close the browser

3. Try Using a Proxy

If your IP has been blocked by the target site, using a proxy can help bypass this:

proxies = {
    'http': 'http://your-proxy-address:port',
    'https': 'https://your-proxy-address:port'
}

# Add proxies to your requests.get call
response = requests.get(url, headers=headers, proxies=proxies, timeout=10)

You can find free proxies online (though they're unreliable) or use a paid proxy service for better results.

4. Check for Anti-Scraping Measures Like Cloudflare

Many sites use Cloudflare or similar services to block bots. If you see a "Checking your browser" page when accessing the site, you might need a library like cfscrape to bypass it:

import cfscrape

scraper = cfscrape.create_scraper()
response = scraper.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

Just note that cfscrape might not work with the latest Cloudflare versions, so you may need to update it regularly.

Final Checks

  • Verify the target URL is correct and accessible in your browser (sometimes typos cause issues)
  • Check your network connection—try accessing other sites to confirm you're online
  • Avoid making too many requests too quickly—add delays between requests with time.sleep(2) to mimic human behavior

内容的提问来源于stack exchange,提问作者tharun amireddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:02:10