使用requests.get爬取网页遭拦截,如何解决该问题?
Hey, I’ve dealt with this exact blocking issue so many times—most websites flag basic requests.get() calls because they don’t look like a real user browsing. Let’s walk through practical fixes that usually work:
1. Add a Proper User-Agent Header
The first thing sites check is your User-Agent string. By default, requests uses a generic one that screams "I’m a script!" Swap it for a real browser’s user agent:
import requests from bs4 import BeautifulSoup # Replace with a current browser user agent (grab yours from your browser's dev tools) headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } req = requests.get('your_target_url', headers=headers) req.encoding='utf-8' html = req.text print(html)
Pro tip: You can find your actual browser’s User-Agent by visiting about:version (Chrome) or checking the Network tab in DevTools.
2. Use a Session to Maintain Cookies
Some sites track sessions with cookies. Using requests.Session() lets you persist cookies across requests, mimicking a real user’s browsing session:
import requests from bs4 import BeautifulSoup session = requests.Session() session.headers.update({ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' }) # First, make a dummy request to grab initial cookies session.get('your_target_url') # Then make your actual request req = session.get('your_target_url') req.encoding='utf-8' html = req.text print(html)
3. Add Delays Between Requests
If you’re scraping multiple pages, hitting the site too fast will get you blocked instantly. Add a random delay to look like a natural user:
import requests from bs4 import BeautifulSoup import time import random headers = {'User-Agent': 'your_user_agent_here'} for url in list_of_urls: req = requests.get(url, headers=headers) # Process your response here... # Wait 2-5 seconds (randomized to avoid pattern detection) time.sleep(random.uniform(2, 5))
4. Rotate Proxy IPs (If All Else Fails)
If the site has strict IP blocking, rotating proxies can help. You can use free proxy lists (unreliable) or paid proxy services for better consistency:
import requests from bs4 import BeautifulSoup proxies = { 'http': 'http://your_proxy_ip:port', 'https': 'https://your_proxy_ip:port' } headers = {'User-Agent': 'your_user_agent_here'} req = requests.get('your_target_url', headers=headers, proxies=proxies)
5. Simulate a Real Browser (Advanced)
For sites with heavy anti-bot measures (like Cloudflare), requests might not cut it. Tools like Selenium or Playwright mimic real browser interactions (including JavaScript rendering):
Selenium example:
from selenium import webdriver from bs4 import BeautifulSoup driver = webdriver.Chrome() driver.get('your_target_url') html = driver.page_source soup = BeautifulSoup(html, 'html.parser') # Process the content... driver.quit()
Start with the first two fixes—they resolve most basic blocking cases. If those don’t work, move to delays or proxies, and save browser simulation for the toughest sites.
内容的提问来源于stack exchange,提问作者박형렬

