You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

VM环境下Scrapy爬取Booking.com返回通用响应问题排查

Troubleshooting Booking.com Scraping Discrepancies Between Local Machine and VM

Let's break down why your VM's Scrapy setup is returning a generic response while your local machine and the Links browser work, and how to fix it:

Key Observations from Your Details

  • Local machine works only after setting the correct User-Agent, which tells us Booking.com uses UA detection as a first line of defense.
  • VM's Links browser works, so your public IP isn't blocked—this points to browser fingerprinting beyond basic headers being the issue.
  • Disabling cookies in the VM didn't fix the problem, so cookies aren't the primary culprit here.

Likely Culprits for the Discrepancy

Booking.com uses sophisticated anti-scraping measures that go beyond User-Agent and robots.txt. Here are the most probable issues:

1. Request Header Format/Order Differences

Looking at your request headers:

  • Local Accept header: text/html,application/xhtml+xml,application/xml;q=0.9,*/*; q=0.8 (note the space after ; before q=0.8)
  • VM Scrapy Accept header: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8 (no space)

Even tiny formatting details like this can trigger anti-scraping filters. Additionally, Scrapy may send headers in a different order than your local browser, which some sites check.

2. TLS/HTTP Client Fingerprinting

Scrapy uses Twisted's HTTP client under the hood, which has a distinct TLS handshake fingerprint compared to real browsers (like Firefox or Links). Booking.com's systems can detect this fingerprint and redirect scrapers to generic pages.

3. URL Parameter Escaping Issue

Your URL contains &#hotelTmpl—this is an HTML-encoded & (&). Real browsers automatically decode this to &#hotelTmpl, but Scrapy might be sending the encoded version, causing Booking.com to ignore the trailing parameter and return a generic page.

Step-by-Step Fixes

1. Exact Header Replication

First, replicate your local browser's headers exactly in Scrapy, including formatting and order. Use the DEFAULT_REQUEST_HEADERS setting to override Scrapy's defaults:

scrapy shell \
--set="ROBOTSTXT_OBEY=False" \
--set="COOKIES_ENABLED=False" \
-s USER_AGENT="Mozilla/5.0 (Android 4.4; Mobile; rv:41.0) Gecko/41.0 Firefox/41.0" \
-s DEFAULT_REQUEST_HEADERS='{"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*; q=0.8", "Accept-Language": "en", "Accept-Encoding": "gzip,deflate"}' \
"https://www.booking.com/hotel/fr/le-transat-bleu.fr.html?aid=304142;label=gen173nr-1FCAEoggJCAlhYSDNiBW5vcmVmaE2IAQGYAQ3CAQp3aW5kb3dzIDEwyAEM2AEB6AEB-AELkgIBeagCAw;sid=746d95cb38d6de7fbb5a878954481e7b;all_sr_blocks=33843609_122840412_1_2_0;checkin=2019-03-17;checkout=2019-03-18;dest_id=-1424668;dest_type=city;dist=0;group_adults=1;group_children=0;hapos=1;highlighted_blocks=33843609_122840412_1_2_0;hpos=1;req_adults=1;req_children=0;room1=A%2C;sb_price_type=total;sr_order=popularity;srepoch=1550502677;srpvid=26936aca347f0334;type=total;ucfs=1&#hotelTmpl"

Notice we replaced &#hotelTmpl with &#hotelTmpl to fix the URL encoding issue.

2. Use a Real Browser with Scrapy Playwright

To bypass TLS fingerprinting, integrate Scrapy with Playwright, which uses real browsers (Chrome, Firefox) to send requests—this matches the fingerprint that Booking.com expects.

  1. Install the package:
pip install scrapy-playwright
  1. Configure Scrapy to use Playwright via command line:
scrapy shell \
--set="ROBOTSTXT_OBEY=False" \
--set="PLAYWRIGHT_ENABLED=True" \
--set="PLAYWRIGHT_LAUNCH_OPTIONS={'headless': True, 'user_agent': 'Mozilla/5.0 (Android 4.4; Mobile; rv:41.0) Gecko/41.0 Firefox/41.0'}" \
"https://www.booking.com/hotel/fr/le-transat-bleu.fr.html?aid=304142;label=gen173nr-1FCAEoggJCAlhYSDNiBW5vcmVmaE2IAQGYAQ3CAQp3aW5kb3dzIDEwyAEM2AEB6AEB-AELkgIBeagCAw;sid=746d95cb38d6de7fbb5a878954481e7b;all_sr_blocks=33843609_122840412_1_2_0;checkin=2019-03-17;checkout=2019-03-18;dest_id=-1424668;dest_type=city;dist=0;group_adults=1;group_children=0;hapos=1;highlighted_blocks=33843609_122840412_1_2_0;hpos=1;req_adults=1;req_children=0;room1=A%2C;sb_price_type=total;sr_order=popularity;srepoch=1550502677;srpvid=26936aca347f0334;type=total;ucfs=1&#hotelTmpl"

3. Validate Requests with Mitmproxy

To confirm exactly what's different between your local and VM requests, use mitmproxy to capture and compare both:

  1. Run mitmproxy on your local machine and configure your browser to use it as a proxy.
  2. Run mitmproxy on your VM and configure Scrapy to use it.
  3. Compare every detail: TLS handshake, request header order, header formatting, and URL parameters.

Final Notes

Booking.com actively updates its anti-scraping measures, so relying solely on header tweaks may not work long-term. Using a real browser via Playwright is the most reliable way to mimic human-like behavior and avoid detection.

内容的提问来源于stack exchange,提问作者Yohan Obadia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:01:07