You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫如何爬取网站全部分类下的所有商品及价格数据

问题排查结论

你的代码只能拿到首页商品,核心原因是你只发起了针对站点根路径https://mrbakeregypt.com/的单次请求,根本没有对各个分类归档页、分类下的分页发起请求,自然拿不到其他位置的商品。除此之外你的代码还有两个显性bug:

  • 你先定义了带请求头的请求对象r,但实际解析时又重新发起了一次不带请求头的requests.get(url)请求,前面配置的headers完全没生效,很容易触发站点反爬
  • 代码里混入了无效的markdown片段`[![enter code here][1]][1]`,直接运行会触发语法错误
修复方案
  • 先从首页导航栏提取所有商品分类的链接,不要只爬首页根路径
  • 每个分类页都做分页遍历:该站点是Shopify架构,分页通过url后拼接?page=页码实现,逐页请求直到页面无商品为止
  • 所有请求统一使用配置好的headers,每次请求后加1-2秒延时,避免被反爬拦截
  • 删除代码中混入的无效片段,修复语法问题
可直接运行的修复后代码
import requests
from bs4 import BeautifulSoup
import pandas as pd
import time

cakes = []
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36'
}

# 提取首页所有商品分类链接
home_url = "https://mrbakeregypt.com/"
home_resp = requests.get(home_url, headers=headers)
home_soup = BeautifulSoup(home_resp.content, "html.parser")
category_urls = []

for nav_link in home_soup.select("nav.site-nav a"):
    link = nav_link.get("href")
    if link and "/collections/" in link:
        # 补全相对路径为完整链接
        if not link.startswith("http"):
            link = f"https://mrbakeregypt.com{link}"
        if link not in category_urls:
            category_urls.append(link)

# 遍历所有分类+分页爬取商品
for cat_url in category_urls:
    page = 1
    while True:
        current_url = f"{cat_url}?page={page}"
        print(f"正在爬取:{current_url}")
        resp = requests.get(current_url, headers=headers)
        # 页面状态异常则停止爬取该分类
        if resp.status_code != 200:
            break
        page_soup = BeautifulSoup(resp.content, "html.parser")
        product_cards = page_soup.find_all("div", class_="grid-view-item product-card")
        # 当前页无商品说明已到最后一页,跳出分页循环
        if not product_cards:
            break
        # 提取当前页商品信息
        for card in product_cards:
            name = card.find("div", class_="h4 grid-view-item__title product-card__title").text.strip()
            price = card.find("div", class_="price__regular").text.strip().replace('\n', '')
            cakes.append({"name": name, "cost": price})
        page += 1
        time.sleep(1.5) # 控制请求频率

# 导出结果
pd.DataFrame(cakes).to_csv("bakery_products.csv", index=False, encoding="utf-8-sig")
print(f"爬取完成,共获取{len(cakes)}件商品,已保存为csv文件")

补充提示:Shopify站点默认开放/products.json接口,支持分页拉取全量商品数据,不需要逐个解析分类页,爬取效率会更高。

内容的提问来源于stack exchange,提问作者Fouad Elmahdy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 00:45:48