使用Python requests.get获取Expedia网页内容返回不完整响应问题
Hey there, let's break down why you're getting incomplete HTML when fetching Expedia's hotel search page with requests.get()—and how to fix it.
1. Incomplete Request Headers Are Triggering Anti-Scraping Checks
Expedia doesn't just look at the User-Agent to validate requests. Browsers send a bunch of additional headers that your current code is missing, which can make the server return truncated content as an anti-scrape measure.
Update your headers to match what a real browser sends:
import requests from lxml import html url = "https://www.expedia.com/Hotel-Search?destination=Maldives&latLong=3.480528%2C73.192127®ionId=109&startDate=04%2F20%2F2018&endDate=04%2F21%2F2018&rooms=1&_xpid=11905%7C1&adults=2" headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/46.0.2490.80 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.expedia.com/', 'DNT': '1', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' } response = requests.get(url, headers=headers) response.encoding = response.apparent_encoding # Ensure correct character encoding html_content = response.text
2. Most Content Is Loaded Dynamically with JavaScript
requests only fetches the initial static HTML sent by the server. Expedia loads almost all hotel listings, prices, and details using JavaScript after the page loads—so that content isn't present in the initial response you're getting.
To get the fully rendered page, use a browser automation tool like Selenium to simulate a real browser that executes JavaScript. Here's a working example:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from lxml import html # Configure Chrome to run in headless mode (no visible window) options = Options() options.add_argument('--headless=new') options.add_argument('user-agent=Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/46.0.2490.80 Safari/537.36') # Initialize the browser and load the page driver = webdriver.Chrome(options=options) driver.get(url) # Wait a few seconds to ensure all dynamic content loads (adjust as needed) driver.implicitly_wait(5) # Get the fully rendered page source page_source = driver.page_source # Parse with lxml as before tree = html.fromstring(page_source) # Don't forget to close the browser when done driver.quit()
3. Slow Down to Avoid Being Blocked
Even with the right headers and tools, sending too many requests too quickly will get your IP blocked by Expedia's anti-scraping systems. Add delays between requests if you're scraping multiple pages:
import time # After each request/page load time.sleep(2) # Wait 2 seconds before the next action
A quick note: Always review Expedia's Terms of Service before scraping—unauthorized scraping can lead to permanent IP blocks or legal action.
内容的提问来源于stack exchange,提问作者Suhail Moideen

