You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python3+BeautifulSoup4爬取家乐福网站价格无结果,求排查

解决家乐福西班牙站爬虫无结果的问题

嘿,作为新手遇到这种情况太正常啦,我帮你捋捋问题出在哪,再给你可行的解决方案~

为什么你的代码拿不到结果?

主要有两个核心原因:

  • 动态内容加载:家乐福西班牙站的搜索结果是通过JavaScript动态渲染出来的。你用requests.get()只能获取到页面的静态HTML骨架,而实际的商品价格、重量这些数据是页面加载完成后才通过AJAX请求加载的,所以静态HTML里根本没有你要找的ebx-result-price__value和ebx-result-quantity类的元素。这也是为什么你的代码在其他静态网站能跑通,但在这里不行的关键原因。
  • 反爬机制拦截:网站可能检测到你的请求来自爬虫(默认的requests请求头没有浏览器标识),返回的内容不完整甚至空白,自然找不到目标元素。

怎么修改代码?

方案1:先尝试添加请求头模拟浏览器

先简单试试给请求加个浏览器的User-Agent,看看能不能拿到完整内容:

import requests
from bs4 import BeautifulSoup

barcodes = ['5449000000996']
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

for barcode in barcodes:
    url = 'https://www.carrefour.es/?q=' + barcode
    html = requests.get(url, headers=headers).content
    bs = BeautifulSoup(html, 'lxml')
    # 先打印页面长度,判断是否拿到完整内容
    print(f"页面内容长度:{len(html)}")
    searchingprice = bs.find_all('strong', {'class':'ebx-result-price__value'})
    print(searchingprice)
    searchingpricerperkg = bs.find_all('span', {'class':'ebx-result__quantity ebx-result-quantity'})
    print(searchingpricerperkg)

如果打印的页面长度很短,或者还是没结果,那基本就是动态加载的问题,得用方案2。

方案2:用Selenium模拟浏览器加载动态内容

Selenium能像真实用户一样打开浏览器,等待页面完全加载(包括JS渲染的内容),这样就能拿到完整的商品数据了。
步骤如下:

  1. 安装Selenium:pip install selenium
  2. 下载对应浏览器的驱动(比如Chrome浏览器要下载ChromeDriver,注意版本要和你的浏览器匹配)
  3. 修改代码:
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

barcodes = ['5449000000996']

# 初始化Chrome浏览器驱动(确保ChromeDriver路径正确,若已配置环境变量可直接写webdriver.Chrome())
driver = webdriver.Chrome()

for barcode in barcodes:
    url = 'https://www.carrefour.es/?q=' + barcode
    driver.get(url)
    
    try:
        # 等待价格元素加载,最多等10秒
        price_element = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CLASS_NAME, 'ebx-result-price__value'))
        )
        # 等待每公斤价格元素加载
        per_kg_element = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CLASS_NAME, 'ebx-result-quantity'))
        )
        
        print(f"商品价格:{price_element.text}")
        print(f"每公斤价格:{per_kg_element.text}")
    except Exception as e:
        print(f"加载元素失败:{e}")
    finally:
        # 循环多个条码时可以暂时不关闭浏览器,方便调试
        pass

# 所有任务完成后关闭浏览器
driver.quit()

小提醒

  • 爬取网站前记得查看网站的robots.txt(比如https://www.carrefour.es/robots.txt),遵守爬取规则,不要频繁发送请求,避免被网站封禁IP。
  • 如果后续遇到元素选择器不对的问题,可以用浏览器的开发者工具(F12)查看真实页面的元素结构,确认class名或者元素路径是否正确。

内容的提问来源于stack exchange,提问作者Pin_Eipol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 10:37:34