You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取CASIO G-SHOCK手表数据时始终返回空列表问题求助

问题

爬取https://www.casio.com/us/watches/gshock/网站上的CASIO手表(价格、名称、型号)信息时,每次都得到空列表。以下是使用BeautifulSoup和requests编写的Python代码及预期输出示例:

原代码

from bs4 import BeautifulSoup
import requests

def get_data(url):
    headers = {
    "User-Agent": "header"
    }
    
    req = requests.get(url="https://www.casio.com/us/watches/gshock/", headers=headers)
    
    # with open("/Users/yerkebulan/Desktop/parser/mylessons/lesson7/watch7/data.html", "w") as file:
    #     file.write(req.text)
    
    soup = BeautifulSoup(req.text, "lxml")
    
    list_wathces = soup.find_all("li", class_="cmp-product_panel_list__item")
    
    for item in list_wathces:
        watch_name = item.find("div", class_ ="cmp-product_panel__code").text
        watch_company = item.find("div", class_ = "cmp-product_panel__code")
        print(watch_company.text)
    
    # with open("/Users/yerkebulan/Desktop/parser/mylessons/lesson7/watch7", "w") as file:
    #     file.
    
def main():
    get_data("https://www.casio.com/us/watches/casio/")

if __name__ == "__main__":
    main()

预期输出示例

[
    {
        "product_article": "GA-700SKE-7AER",
        "product_url": "https://shop.casio.ru/catalog/g-shock-youth/ga-700ske-7aer/",
        "product_price": "12 390"
    },
    {
        "product_article": "GA-2100SKE-7AER",
        "product_url": "https://shop.casio.ru/catalog/g-shock-youth/ga-2100ske-7aer/",
        "product_price": "11 790"
    }
]
解决方案

直接梳理核心问题及修复步骤:

  1. URL匹配错误:main()函数调用的是https://www.casio.com/us/watches/casio/,但目标爬取页面是G-Shock专区,URL不对应直接导致爬错页面,自然拿不到目标元素。
  2. User-Agent无效:headers里的User-Agent设为"header",会被网站判定为异常请求,返回的静态HTML不包含产品列表。需替换为真实浏览器UA,比如Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36。
  3. 页面动态加载:Casio官网产品列表通过JavaScript动态渲染,requests只能获取静态HTML,无法拿到JS加载后的内容。必须用浏览器模拟工具(如Selenium、Playwright)加载完整页面。
  4. 元素定位失效:原代码中的cmp-product_panel_list__item、cmp-product_panel__code等类名可能已过时,需通过浏览器开发者工具重新确认当前页面的DOM结构,调整定位器。

修正后的代码(基于Selenium)

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time

def get_data(url):
    # 配置无头浏览器选项
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")
    chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    
    driver = webdriver.Chrome(options=chrome_options)
    driver.get(url)
    time.sleep(3)  # 等待动态内容加载完成
    
    # 获取渲染后的页面源码
    soup = BeautifulSoup(driver.page_source, "lxml")
    driver.quit()
    
    # 注意:以下类名需根据当前页面实际DOM调整,此处为示例
    list_watches = soup.find_all("li", class_="product-item")
    result = []
    
    for item in list_watches:
        product_article = item.find("div", class_="product-code").text.strip() if item.find("div", class_="product-code") else "未知型号"
        product_url = item.find("a")["href"] if item.find("a") else "无链接"
        product_price = item.find("span", class_="price").text.strip() if item.find("span", class_="price") else "无价格"
        
        result.append({
            "product_article": product_article,
            "product_url": product_url,
            "product_price": product_price
        })
    
    return result

def main():
    watch_data = get_data("https://www.casio.com/us/watches/gshock/")
    print(watch_data)

if __name__ == "__main__":
    main()

注意事项

  • 运行前需安装Selenium及对应浏览器驱动(如ChromeDriver)。
  • 网站DOM结构可能随时更新,使用前务必通过浏览器开发者工具确认元素的最新类名或选择器。
  • 可替换固定sleep为Selenium的显式等待,提升爬取稳定性。

内容的提问来源于stack exchange,提问作者Yerkebulan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 22:47:34