You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests.get爬取网页遭拦截,如何解决该问题?

Fixing "Sorry, You Have Been Blocked" with Python Requests

Hey, I’ve dealt with this exact blocking issue so many times—most websites flag basic requests.get() calls because they don’t look like a real user browsing. Let’s walk through practical fixes that usually work:

1. Add a Proper User-Agent Header

The first thing sites check is your User-Agent string. By default, requests uses a generic one that screams "I’m a script!" Swap it for a real browser’s user agent:

import requests
from bs4 import BeautifulSoup

# Replace with a current browser user agent (grab yours from your browser's dev tools)
headers = {
    'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

req = requests.get('your_target_url', headers=headers)
req.encoding='utf-8'
html = req.text
print(html)

Pro tip: You can find your actual browser’s User-Agent by visiting about:version (Chrome) or checking the Network tab in DevTools.

2. Use a Session to Maintain Cookies

Some sites track sessions with cookies. Using requests.Session() lets you persist cookies across requests, mimicking a real user’s browsing session:

import requests
from bs4 import BeautifulSoup

session = requests.Session()
session.headers.update({
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
})

# First, make a dummy request to grab initial cookies
session.get('your_target_url')
# Then make your actual request
req = session.get('your_target_url')
req.encoding='utf-8'
html = req.text
print(html)

3. Add Delays Between Requests

If you’re scraping multiple pages, hitting the site too fast will get you blocked instantly. Add a random delay to look like a natural user:

import requests
from bs4 import BeautifulSoup
import time
import random

headers = {'User-Agent': 'your_user_agent_here'}

for url in list_of_urls:
    req = requests.get(url, headers=headers)
    # Process your response here...
    # Wait 2-5 seconds (randomized to avoid pattern detection)
    time.sleep(random.uniform(2, 5))

4. Rotate Proxy IPs (If All Else Fails)

If the site has strict IP blocking, rotating proxies can help. You can use free proxy lists (unreliable) or paid proxy services for better consistency:

import requests
from bs4 import BeautifulSoup

proxies = {
    'http': 'http://your_proxy_ip:port',
    'https': 'https://your_proxy_ip:port'
}

headers = {'User-Agent': 'your_user_agent_here'}
req = requests.get('your_target_url', headers=headers, proxies=proxies)

5. Simulate a Real Browser (Advanced)

For sites with heavy anti-bot measures (like Cloudflare), requests might not cut it. Tools like Selenium or Playwright mimic real browser interactions (including JavaScript rendering):

Selenium example:

from selenium import webdriver
from bs4 import BeautifulSoup

driver = webdriver.Chrome()
driver.get('your_target_url')
html = driver.page_source
soup = BeautifulSoup(html, 'html.parser')
# Process the content...
driver.quit()

Start with the first two fixes—they resolve most basic blocking cases. If those don’t work, move to delays or proxies, and save browser simulation for the toughest sites.

内容的提问来源于stack exchange,提问作者박형렬

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:39:12