Python Request遭Cloudflare拦截:设置Session仍返回检查页面原因
Ah, I’ve dealt with this exact headache more times than I can count—Cloudflare’s bot detection is stubborn, and just using a requests.Session() doesn’t automatically bypass it. Let’s break down why this happens and how you can fix it.
Common Reasons the Session Isn’t Working
Session only handles cookies, not full browser behavior
Cloudflare doesn’t just check for cookies. It analyzes a ton of signals to tell bots apart from real users: valid user-agent strings, JavaScript execution, browser fingerprint details (like screen resolution, canvas rendering), and even how you interact with the page (mouse movements, click patterns). A basicSessiononly persists cookies between requests—it doesn’t mimic any of these real-browser traits, so Cloudflare flags you immediately.Missing or invalid request headers
Real browsers send a whole suite of headers with every request, and Cloudflare expects to see them. If your Session is using a generic or outdatedUser-Agent, or missing headers likeAccept-Language,Accept-Encoding, orReferer, Cloudflare will know something’s off.JavaScript challenge isn’t being resolved
That "Checking your browser" page is often a JavaScript-based check.requests.Session()can’t execute JavaScript—it just fetches raw HTML. So Cloudflare never gets the confirmation that you’re a real browser that can run its scripts, and it keeps blocking you.
Fixes to Get Past the Check
Use a headless browser tool
Tools like Selenium, Playwright, or Pyppeteer mimic real browsers completely—they run JavaScript, handle cookies automatically, and generate valid browser fingerprints. You can keep a single browser instance open (like a session) to maintain your state across requests. This is the most reliable way to bypass Cloudflare’s checks for most cases.Add realistic headers to your Session
If you want to stick withrequests.Session(), update your headers to match what a real browser sends. For example:import requests session = requests.Session() session.headers.update({ 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 14_0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9', 'Accept-Encoding': 'gzip, deflate, br', 'Referer': 'https://the-target-website.com/', 'DNT': '1' })Just make sure your user-agent is a current, widely used one—avoid generic strings like "python-requests/2.31.0".
Manually pass Cloudflare cookies (temporary fix)
You can visit the target site in a real browser, solve the Cloudflare check, then copy thecf_clearancecookie and your browser’s user-agent. Add both to your Session:session.cookies.set('cf_clearance', 'your-cf-clearance-value') session.headers['User-Agent'] = 'your-browser-user-agent'Note: This cookie is tied to your IP and user-agent, so if either changes, it’ll stop working. It also expires after a few hours.
Use a specialized library
Libraries likecloudscraperare built to bypass Cloudflare’s detection by mimicking browser JavaScript execution. It works just likerequests.Session(), so it’s easy to drop into your code:import cloudscraper scraper = cloudscraper.create_scraper() # Creates a session-like object response = scraper.get('https://target-website.com') print(response.text)
A quick reminder: Always make sure you’re following the website’s terms of service when automating access—scraping or botting might be against their rules.
内容的提问来源于stack exchange,提问作者Pthomas

