You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

浏览器检查器与Request获取的HTML内容不一致的原因及解决方法

解决网页爬取内容与浏览器不一致的问题

问题背景

要爬取https://www.onestopwineshop.com/collection/type/red-wines的红酒数据,用requests+BeautifulSoup拿到的HTML和浏览器里的内容对不上,怀疑是网站识别出非浏览器访问限制了数据。试了Selenium也没解决问题。

原Requests代码

import requests
from bs4 import BeautifulSoup
url = "https://www.onestopwineshop.com/collection/type/red-wines"
response = requests.get(url)
#print(response.text)
soup = BeautifulSoup(response.content,'lxml')

原Selenium代码

from selenium import webdriver
import time
path = "C:\Program Files (x86)\chromedriver.exe"
# start web browser
browser=webdriver.Chrome(path)
#navigate to the page
url = "https://www.onestopwineshop.com/collection/type/red-wines"
browser.get(url)
# sleep the required amount to let the page load
time.sleep(3)
# get source code
html = browser.page_source
# close web browser
browser.close()

开发者工具页面截图

开发者工具页面截图

可行解决方法

1. 给Requests补全请求头,模拟真实浏览器

网站靠请求头识别非浏览器,把User-Agent、Accept这些关键字段补全:

import requests
from bs4 import BeautifulSoup

url = "https://www.onestopwineshop.com/collection/type/red-wines"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8",
    "Accept-Language": "zh-CN,zh;q=0.8,en-US;q=0.5,en;q=0.3",
    "Referer": "https://www.onestopwineshop.com/",
    "Connection": "keep-alive"
}

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.content, 'lxml')

2. 升级Selenium,绕过网站的自动化检测

普通Selenium容易被识破,给Chrome加些配置,模拟真人浏览器:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
import time

options = Options()
# 关掉自动化控制提示
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)
# 模拟真实浏览器的UA
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
# 可选:禁用图片加载,提速
options.add_argument("--blink-settings=imagesEnabled=false")

# 启动浏览器
browser = webdriver.Chrome(options=options)
browser.get("https://www.onestopwineshop.com/collection/type/red-wines")

# 别用固定sleep,等商品元素加载出来再拿源码
try:
    # 这里的CLASS_NAME要换成页面实际的商品类名,自己去浏览器里查
    WebDriverWait(browser, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "product-item"))
    )
except:
    pass

html = browser.page_source
browser.close()

3. 用undetected-chromedriver绕Cloudflare验证

从截图看大概率是Cloudflare的人机验证,普通Selenium搞不定,试试专门的库:
先安装:pip install undetected-chromedriver
然后用这个代码:

import undetected_chromedriver as uc
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC

driver = uc.Chrome()
driver.get("https://www.onestopwineshop.com/collection/type/red-wines")

# 等内容加载好
WebDriverWait(driver, 15).until(
    EC.presence_of_element_located((By.CLASS_NAME, "product-item"))
)

html = driver.page_source
driver.close()

4. 直接抓API接口更省心

打开浏览器开发者工具的「网络」标签,刷新页面,找XHR/Fetch请求,电商网站一般会用API加载商品数据。找到带products或者red-wines的接口,复制请求地址和参数,直接用requests调用,比爬页面效率高多了。

内容的提问来源于stack exchange,提问作者DJ-coding

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 10:36:04