浏览器检查器与Request获取的HTML内容不一致的原因及解决方法
解决网页爬取内容与浏览器不一致的问题
问题背景
要爬取https://www.onestopwineshop.com/collection/type/red-wines的红酒数据,用requests+BeautifulSoup拿到的HTML和浏览器里的内容对不上,怀疑是网站识别出非浏览器访问限制了数据。试了Selenium也没解决问题。
原Requests代码
import requests from bs4 import BeautifulSoup url = "https://www.onestopwineshop.com/collection/type/red-wines" response = requests.get(url) #print(response.text) soup = BeautifulSoup(response.content,'lxml')
原Selenium代码
from selenium import webdriver import time path = "C:\Program Files (x86)\chromedriver.exe" # start web browser browser=webdriver.Chrome(path) #navigate to the page url = "https://www.onestopwineshop.com/collection/type/red-wines" browser.get(url) # sleep the required amount to let the page load time.sleep(3) # get source code html = browser.page_source # close web browser browser.close()
开发者工具页面截图

可行解决方法
1. 给Requests补全请求头,模拟真实浏览器
网站靠请求头识别非浏览器,把User-Agent、Accept这些关键字段补全:
import requests from bs4 import BeautifulSoup url = "https://www.onestopwineshop.com/collection/type/red-wines" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8", "Accept-Language": "zh-CN,zh;q=0.8,en-US;q=0.5,en;q=0.3", "Referer": "https://www.onestopwineshop.com/", "Connection": "keep-alive" } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.content, 'lxml')
2. 升级Selenium,绕过网站的自动化检测
普通Selenium容易被识破,给Chrome加些配置,模拟真人浏览器:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC import time options = Options() # 关掉自动化控制提示 options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) # 模拟真实浏览器的UA options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") # 可选:禁用图片加载,提速 options.add_argument("--blink-settings=imagesEnabled=false") # 启动浏览器 browser = webdriver.Chrome(options=options) browser.get("https://www.onestopwineshop.com/collection/type/red-wines") # 别用固定sleep,等商品元素加载出来再拿源码 try: # 这里的CLASS_NAME要换成页面实际的商品类名,自己去浏览器里查 WebDriverWait(browser, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "product-item")) ) except: pass html = browser.page_source browser.close()
3. 用undetected-chromedriver绕Cloudflare验证
从截图看大概率是Cloudflare的人机验证,普通Selenium搞不定,试试专门的库:
先安装:pip install undetected-chromedriver
然后用这个代码:
import undetected_chromedriver as uc from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC driver = uc.Chrome() driver.get("https://www.onestopwineshop.com/collection/type/red-wines") # 等内容加载好 WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.CLASS_NAME, "product-item")) ) html = driver.page_source driver.close()
4. 直接抓API接口更省心
打开浏览器开发者工具的「网络」标签,刷新页面,找XHR/Fetch请求,电商网站一般会用API加载商品数据。找到带products或者red-wines的接口,复制请求地址和参数,直接用requests调用,比爬页面效率高多了。
内容的提问来源于stack exchange,提问作者DJ-coding
相关产品推荐
相关产品推荐

