Selenium爬取e-dreams酒店价格偶发失败问题排查
E-dreams酒店爬虫价格数据时有时无问题解决
问题背景
使用Python Selenium开发e-dreams酒店数据爬取程序,目标抓取酒店名称、当前价格并存储为数据集。程序执行时仅价格字段出现异常:有时能正常存储,有时无法获取,其他数据无异常,怀疑是反爬机制或动态加载导致。
原代码
import selenium import requests import os from selenium import webdriver from bs4 import BeautifulSoup from selenium.common.exceptions import WebDriverException from selenium.webdriver.chrome.options import Options import pandas as pd from datetime import date import time from selenium.webdriver.common.by import By opts=Options() opts.add_argument("user-agent=Mozilla/5.0 (X11; CrOS x86_64 8172.45.0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/51.0.2704.64 Safari/537.36") driver=webdriver.Chrome("C:\Program Files (x86)\chromedriver.exe", chrome_options=opts) link='https://hotels.edreams.com/searchresults.html?aid=350435&checkin=2022-11-30&checkout=2022-12-01&ctoken=JfaDoWzscPcWaAMp8gw419frpsZ1zq0hhSN70OhXDJn3rXElJMlalObVxI1Y8fIJ2USoEpG_paizTh6YXrsxIRE3Vd8VdsRuWcL1BjSLkEWEGCS9n_miFb3LqswF5npXnOo2kWB_BcwuzzFUqUDsGIVUkQqBp5I3uQHxjE2ixEreVopkod_PsP7q1AXBD02SDe3zGSBdgKCpmWS1g2aiVThU6mWfLql3URPKRhweuZPevAra2DVMGnPcwrzLhlfPY6lE1uoLQWgOkZrdiu64IkegbctyGWuZzd3JpoBIZHWr2oti7CZOyDYUPnknJcG0bDC6OjSFQ6J1ZW6on3BjILPK9Wq0E_JKqrMs4IH0IRF96Txf4E_MtWNi34JLK5n5fQlqcc5UynwtQszkNFF0W7t6GKTvtAfQfW1Srwj3Rn8894DbBLnbtDfOOgs&dest_id=-2601889&dest_type=city&fp_referrer_aid=308918&group_adults=1&group_children=0&label=edr-link-com-sb-conf-pc-of&lang=en-gb&no_rooms=1&selected_currency=EUR&si=ai%2Cco%2Cci%2Cre%2Cdi&sp_plprd=UmFuZG9tSVYkc2RlIyh9YVXcKaaJl1ClKWK-8iYFdtMHKztuOGDCrZfdqiRGdeskuj-OEKqKFdayZs2UNKGuqQQHOdEClFirVT_0eoZ8amf7u6qjvNJ_hAHMLhfjuMoQH81__grdeE3tynZUX-P9WQ&ss=London%2C%20Greater%20London%2C%20United%20Kingdom&submit=Search%20hotels&utm_campaign=%28organic%29&utm_medium=organic&utm_source=google&utm_term=edreams&' page=BeautifulSoup(driver.page_source,'html.parser') hotel_name=[] price_now=[] for i in page.findAll('div',attrs={'class':'d20f4628d0'}): title=i.find("div",attrs={'fcab3ed991 a23c043802'}) if title: hotel_name.append(title.text) else: hotel_name.append('') price=i.find("span",attrs={'class':'fcab3ed991 fbd1d3018c e729ed5ab6'}) if price: price_now.append(price.text) else: price_now.append('') df=pd.DataFrame({'Name':hotel_name, 'Price':price_now}) driver.quit()
价格元素信息
价格对应的HTML元素类为 fcab3ed991 fbd1d3018c e729ed5ab6
解决方案
1. 等待动态元素加载完成
价格字段大概率是通过JS动态渲染的,直接获取页面源码可能在元素未加载完成时就执行了抓取。改用Selenium显式等待,确保元素加载后再操作:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 打开链接 driver.get(link) # 显式等待价格元素加载,最长等待10秒 WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "span.fcab3ed991.fbd1d3018c.e729ed5ab6")) ) # 再获取页面源码 page = BeautifulSoup(driver.page_source, 'html.parser')
2. 升级反爬规避策略
- 更新User-Agent:原代码使用的Chrome 51版本过于老旧,容易被反爬识别,替换为现代浏览器UA:
opts.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
- 隐藏自动化特征:添加参数避免Selenium被网站检测:
opts.add_argument("--disable-blink-features=AutomationControlled") opts.add_experimental_option("excludeSwitches", ["enable-automation"]) opts.add_experimental_option('useAutomationExtension', False)
3. 直接用Selenium定位元素
跳过BeautifulSoup解析静态源码的步骤,直接用Selenium定位元素,更适配动态页面:
hotel_name = [] price_now = [] # 定位所有酒店容器 hotel_containers = driver.find_elements(By.CSS_SELECTOR, "div.d20f4628d0") for container in hotel_containers: # 获取酒店名称 try: title = container.find_element(By.CSS_SELECTOR, "div.fcab3ed991.a23c043802").text except: title = "" hotel_name.append(title) # 获取价格 try: price = container.find_element(By.CSS_SELECTOR, "span.fcab3ed991.fbd1d3018c.e729ed5ab6").text except: price = "" price_now.append(price)
4. 添加随机延迟模拟人类操作
在关键步骤添加随机延迟,避免请求频率过高触发反爬:
import random # 页面加载后添加2-5秒随机延迟 time.sleep(random.uniform(2, 5)) # 循环抓取时添加0.5-1.5秒随机延迟 for container in hotel_containers: # ... 抓取代码 ... time.sleep(random.uniform(0.5, 1.5))
内容的提问来源于stack exchange,提问作者Stackoverflow Stackoverflow
相关产品推荐
相关产品推荐

