You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Sklavenitis电商仅获首个商品数据问题求助(教育用途)

解决Sklavenitis网站饼干类商品爬取不全的问题

问题背景

作为编程新手,为实战项目爬取Sklavenitis网站的饼干类商品数据,当前代码仅能获取第一个商品的数据,此前曾成功爬取前6个商品但丢失对应代码,代码仅用于个人学习。

原代码如下:

import requests
from bs4 import BeautifulSoup
import os
from datetime import datetime

# Define the URLs you want to scrape
urls = ["https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/",
        "https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/?pg=2",
        "https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/?pg=3"]

for url in urls:
    # Make a request
    response = requests.get(url)

    # Parse the HTML content
    soup = BeautifulSoup(response.content, 'html.parser')

    # Find all the products on the page
    products = soup.find_all('div', class_='product')

    # Define the file path
    timestamp = datetime.now().strftime("%Y-%m-%d %H-%M-%S")
    filename = os.path.join(os.path.expanduser('~'), 'Desktop', f'{timestamp}.txt')

    # Write the product information to a file on the desktop
    with open(filename, "a") as f:  # Use 'a' instead of 'w' to append to the file instead of overwriting it
        # Add timestamp header for the first URL
        if url == urls[0]:
            f.write("Time, Product Title, Price, Deleted Price\n")

        # Loop over each product and extract the information
        for product in products:
            product_title = product.find('h4', class_='product__title').text.strip()
            deleted_price = product.find('div', class_='deleted__price').text.strip()
            current_price = product.find('div', class_='price')['data-price']

            # Add timestamp column
            f.write(f"{datetime.now().strftime('%Y-%m-%d %H:%M:%S')}, {product_title}, {current_price}, {deleted_price}\n")

    print(f"Product information from {url} saved to desktop!")

问题排查

原代码存在三个关键问题:

  • 缺少请求头:直接用requests.get()请求会被网站识别为非浏览器请求,可能返回不完整的页面内容
  • 未处理缺失元素:假设所有商品都有deleted__price元素,一旦某个商品无此元素,代码会抛出AttributeError并中断执行,导致仅爬取到第一个商品就停止
  • 文件生成逻辑错误:每次循环URL都会生成新的时间戳文件,三个URL对应三个不同文件,数据分散不易查看

修复后的代码

import requests
from bs4 import BeautifulSoup
import os
from datetime import datetime

# 定义请求头,模拟浏览器访问
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

urls = [
    "https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/",
    "https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/?pg=2",
    "https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/?pg=3"
]

# 提前生成唯一文件名,所有页面数据写入同一个文件
timestamp = datetime.now().strftime("%Y-%m-%d %H-%M-%S")
filename = os.path.join(os.path.expanduser('~'), 'Desktop', f'{timestamp}.txt')

# 先写入表头
with open(filename, "w") as f:
    f.write("Time, Product Title, Price, Deleted Price\n")

for url in urls:
    # 添加请求头发送请求
    response = requests.get(url, headers=headers)
    # 确保请求成功
    if response.status_code != 200:
        print(f"请求{url}失败,状态码:{response.status_code}")
        continue

    soup = BeautifulSoup(response.content, 'html.parser')
    products = soup.find_all('div', class_='product')

    print(f"从{url}获取到{len(products)}个商品")

    with open(filename, "a", encoding='utf-8') as f:
        for product in products:
            # 提取商品标题
            title_elem = product.find('h4', class_='product__title')
            product_title = title_elem.text.strip() if title_elem else "无标题"
            
            # 提取当前价格
            price_elem = product.find('div', class_='price')
            current_price = price_elem['data-price'] if price_elem else "无价格"
            
            # 提取原价(处理缺失情况)
            deleted_price_elem = product.find('div', class_='deleted__price')
            deleted_price = deleted_price_elem.text.strip() if deleted_price_elem else "无原价"

            # 写入数据
            f.write(f"{datetime.now().strftime('%Y-%m-%d %H:%M:%S')}, {product_title}, {current_price}, {deleted_price}\n")

    print(f"{url}的商品信息已追加到文件!")

print(f"所有数据已保存到桌面文件:{filename}")

关键修改说明

  • 添加User-Agent请求头,避免被网站反爬拦截
  • 对每个元素的提取增加判断,防止因元素缺失导致代码中断
  • 将文件名生成移到循环外,所有页面数据写入同一个文件,方便查看
  • 增加请求状态码检查,便于排查请求失败问题
  • 写入文件时指定utf-8编码,避免中文/特殊字符乱码

内容的提问来源于stack exchange,提问作者Koutsa Koutsa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 04:18:15