爬取CASIO G-SHOCK手表数据时始终返回空列表问题求助
问题
爬取https://www.casio.com/us/watches/gshock/网站上的CASIO手表(价格、名称、型号)信息时,每次都得到空列表。以下是使用BeautifulSoup和requests编写的Python代码及预期输出示例:
原代码
from bs4 import BeautifulSoup import requests def get_data(url): headers = { "User-Agent": "header" } req = requests.get(url="https://www.casio.com/us/watches/gshock/", headers=headers) # with open("/Users/yerkebulan/Desktop/parser/mylessons/lesson7/watch7/data.html", "w") as file: # file.write(req.text) soup = BeautifulSoup(req.text, "lxml") list_wathces = soup.find_all("li", class_="cmp-product_panel_list__item") for item in list_wathces: watch_name = item.find("div", class_ ="cmp-product_panel__code").text watch_company = item.find("div", class_ = "cmp-product_panel__code") print(watch_company.text) # with open("/Users/yerkebulan/Desktop/parser/mylessons/lesson7/watch7", "w") as file: # file. def main(): get_data("https://www.casio.com/us/watches/casio/") if __name__ == "__main__": main()
预期输出示例
[ { "product_article": "GA-700SKE-7AER", "product_url": "https://shop.casio.ru/catalog/g-shock-youth/ga-700ske-7aer/", "product_price": "12 390" }, { "product_article": "GA-2100SKE-7AER", "product_url": "https://shop.casio.ru/catalog/g-shock-youth/ga-2100ske-7aer/", "product_price": "11 790" } ]
解决方案
直接梳理核心问题及修复步骤:
- URL匹配错误:
main()函数调用的是https://www.casio.com/us/watches/casio/,但目标爬取页面是G-Shock专区,URL不对应直接导致爬错页面,自然拿不到目标元素。 - User-Agent无效:headers里的
User-Agent设为"header",会被网站判定为异常请求,返回的静态HTML不包含产品列表。需替换为真实浏览器UA,比如Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36。 - 页面动态加载:Casio官网产品列表通过JavaScript动态渲染,
requests只能获取静态HTML,无法拿到JS加载后的内容。必须用浏览器模拟工具(如Selenium、Playwright)加载完整页面。 - 元素定位失效:原代码中的
cmp-product_panel_list__item、cmp-product_panel__code等类名可能已过时,需通过浏览器开发者工具重新确认当前页面的DOM结构,调整定位器。
修正后的代码(基于Selenium)
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import time def get_data(url): # 配置无头浏览器选项 chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) driver.get(url) time.sleep(3) # 等待动态内容加载完成 # 获取渲染后的页面源码 soup = BeautifulSoup(driver.page_source, "lxml") driver.quit() # 注意:以下类名需根据当前页面实际DOM调整,此处为示例 list_watches = soup.find_all("li", class_="product-item") result = [] for item in list_watches: product_article = item.find("div", class_="product-code").text.strip() if item.find("div", class_="product-code") else "未知型号" product_url = item.find("a")["href"] if item.find("a") else "无链接" product_price = item.find("span", class_="price").text.strip() if item.find("span", class_="price") else "无价格" result.append({ "product_article": product_article, "product_url": product_url, "product_price": product_price }) return result def main(): watch_data = get_data("https://www.casio.com/us/watches/gshock/") print(watch_data) if __name__ == "__main__": main()
注意事项
- 运行前需安装Selenium及对应浏览器驱动(如ChromeDriver)。
- 网站DOM结构可能随时更新,使用前务必通过浏览器开发者工具确认元素的最新类名或选择器。
- 可替换固定
sleep为Selenium的显式等待,提升爬取稳定性。
内容的提问来源于stack exchange,提问作者Yerkebulan
相关产品推荐
相关产品推荐

