VM环境下Scrapy爬取Booking.com返回通用响应问题排查
Let's break down why your VM's Scrapy setup is returning a generic response while your local machine and the Links browser work, and how to fix it:
Key Observations from Your Details
- Local machine works only after setting the correct User-Agent, which tells us Booking.com uses UA detection as a first line of defense.
- VM's Links browser works, so your public IP isn't blocked—this points to browser fingerprinting beyond basic headers being the issue.
- Disabling cookies in the VM didn't fix the problem, so cookies aren't the primary culprit here.
Likely Culprits for the Discrepancy
Booking.com uses sophisticated anti-scraping measures that go beyond User-Agent and robots.txt. Here are the most probable issues:
1. Request Header Format/Order Differences
Looking at your request headers:
- Local
Acceptheader:text/html,application/xhtml+xml,application/xml;q=0.9,*/*; q=0.8(note the space after;beforeq=0.8) - VM Scrapy
Acceptheader:text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8(no space)
Even tiny formatting details like this can trigger anti-scraping filters. Additionally, Scrapy may send headers in a different order than your local browser, which some sites check.
2. TLS/HTTP Client Fingerprinting
Scrapy uses Twisted's HTTP client under the hood, which has a distinct TLS handshake fingerprint compared to real browsers (like Firefox or Links). Booking.com's systems can detect this fingerprint and redirect scrapers to generic pages.
3. URL Parameter Escaping Issue
Your URL contains &#hotelTmpl—this is an HTML-encoded & (&). Real browsers automatically decode this to &#hotelTmpl, but Scrapy might be sending the encoded version, causing Booking.com to ignore the trailing parameter and return a generic page.
Step-by-Step Fixes
1. Exact Header Replication
First, replicate your local browser's headers exactly in Scrapy, including formatting and order. Use the DEFAULT_REQUEST_HEADERS setting to override Scrapy's defaults:
scrapy shell \ --set="ROBOTSTXT_OBEY=False" \ --set="COOKIES_ENABLED=False" \ -s USER_AGENT="Mozilla/5.0 (Android 4.4; Mobile; rv:41.0) Gecko/41.0 Firefox/41.0" \ -s DEFAULT_REQUEST_HEADERS='{"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*; q=0.8", "Accept-Language": "en", "Accept-Encoding": "gzip,deflate"}' \ "https://www.booking.com/hotel/fr/le-transat-bleu.fr.html?aid=304142;label=gen173nr-1FCAEoggJCAlhYSDNiBW5vcmVmaE2IAQGYAQ3CAQp3aW5kb3dzIDEwyAEM2AEB6AEB-AELkgIBeagCAw;sid=746d95cb38d6de7fbb5a878954481e7b;all_sr_blocks=33843609_122840412_1_2_0;checkin=2019-03-17;checkout=2019-03-18;dest_id=-1424668;dest_type=city;dist=0;group_adults=1;group_children=0;hapos=1;highlighted_blocks=33843609_122840412_1_2_0;hpos=1;req_adults=1;req_children=0;room1=A%2C;sb_price_type=total;sr_order=popularity;srepoch=1550502677;srpvid=26936aca347f0334;type=total;ucfs=1&#hotelTmpl"
Notice we replaced &#hotelTmpl with &#hotelTmpl to fix the URL encoding issue.
2. Use a Real Browser with Scrapy Playwright
To bypass TLS fingerprinting, integrate Scrapy with Playwright, which uses real browsers (Chrome, Firefox) to send requests—this matches the fingerprint that Booking.com expects.
- Install the package:
pip install scrapy-playwright
- Configure Scrapy to use Playwright via command line:
scrapy shell \ --set="ROBOTSTXT_OBEY=False" \ --set="PLAYWRIGHT_ENABLED=True" \ --set="PLAYWRIGHT_LAUNCH_OPTIONS={'headless': True, 'user_agent': 'Mozilla/5.0 (Android 4.4; Mobile; rv:41.0) Gecko/41.0 Firefox/41.0'}" \ "https://www.booking.com/hotel/fr/le-transat-bleu.fr.html?aid=304142;label=gen173nr-1FCAEoggJCAlhYSDNiBW5vcmVmaE2IAQGYAQ3CAQp3aW5kb3dzIDEwyAEM2AEB6AEB-AELkgIBeagCAw;sid=746d95cb38d6de7fbb5a878954481e7b;all_sr_blocks=33843609_122840412_1_2_0;checkin=2019-03-17;checkout=2019-03-18;dest_id=-1424668;dest_type=city;dist=0;group_adults=1;group_children=0;hapos=1;highlighted_blocks=33843609_122840412_1_2_0;hpos=1;req_adults=1;req_children=0;room1=A%2C;sb_price_type=total;sr_order=popularity;srepoch=1550502677;srpvid=26936aca347f0334;type=total;ucfs=1&#hotelTmpl"
3. Validate Requests with Mitmproxy
To confirm exactly what's different between your local and VM requests, use mitmproxy to capture and compare both:
- Run mitmproxy on your local machine and configure your browser to use it as a proxy.
- Run mitmproxy on your VM and configure Scrapy to use it.
- Compare every detail: TLS handshake, request header order, header formatting, and URL parameters.
Final Notes
Booking.com actively updates its anti-scraping measures, so relying solely on header tweaks may not work long-term. Using a real browser via Playwright is the most reliable way to mimic human-like behavior and avoid detection.
内容的提问来源于stack exchange,提问作者Yohan Obadia

