Python爬取Sklavenitis电商仅获首个商品数据问题求助(教育用途)
解决Sklavenitis网站饼干类商品爬取不全的问题
问题背景
作为编程新手,为实战项目爬取Sklavenitis网站的饼干类商品数据,当前代码仅能获取第一个商品的数据,此前曾成功爬取前6个商品但丢失对应代码,代码仅用于个人学习。
原代码如下:
import requests from bs4 import BeautifulSoup import os from datetime import datetime # Define the URLs you want to scrape urls = ["https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/", "https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/?pg=2", "https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/?pg=3"] for url in urls: # Make a request response = requests.get(url) # Parse the HTML content soup = BeautifulSoup(response.content, 'html.parser') # Find all the products on the page products = soup.find_all('div', class_='product') # Define the file path timestamp = datetime.now().strftime("%Y-%m-%d %H-%M-%S") filename = os.path.join(os.path.expanduser('~'), 'Desktop', f'{timestamp}.txt') # Write the product information to a file on the desktop with open(filename, "a") as f: # Use 'a' instead of 'w' to append to the file instead of overwriting it # Add timestamp header for the first URL if url == urls[0]: f.write("Time, Product Title, Price, Deleted Price\n") # Loop over each product and extract the information for product in products: product_title = product.find('h4', class_='product__title').text.strip() deleted_price = product.find('div', class_='deleted__price').text.strip() current_price = product.find('div', class_='price')['data-price'] # Add timestamp column f.write(f"{datetime.now().strftime('%Y-%m-%d %H:%M:%S')}, {product_title}, {current_price}, {deleted_price}\n") print(f"Product information from {url} saved to desktop!")
问题排查
原代码存在三个关键问题:
- 缺少请求头:直接用
requests.get()请求会被网站识别为非浏览器请求,可能返回不完整的页面内容 - 未处理缺失元素:假设所有商品都有
deleted__price元素,一旦某个商品无此元素,代码会抛出AttributeError并中断执行,导致仅爬取到第一个商品就停止 - 文件生成逻辑错误:每次循环URL都会生成新的时间戳文件,三个URL对应三个不同文件,数据分散不易查看
修复后的代码
import requests from bs4 import BeautifulSoup import os from datetime import datetime # 定义请求头,模拟浏览器访问 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } urls = [ "https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/", "https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/?pg=2", "https://www.sklavenitis.gr/mpiskota-sokolates-zacharodi/mpiskota/?pg=3" ] # 提前生成唯一文件名,所有页面数据写入同一个文件 timestamp = datetime.now().strftime("%Y-%m-%d %H-%M-%S") filename = os.path.join(os.path.expanduser('~'), 'Desktop', f'{timestamp}.txt') # 先写入表头 with open(filename, "w") as f: f.write("Time, Product Title, Price, Deleted Price\n") for url in urls: # 添加请求头发送请求 response = requests.get(url, headers=headers) # 确保请求成功 if response.status_code != 200: print(f"请求{url}失败,状态码:{response.status_code}") continue soup = BeautifulSoup(response.content, 'html.parser') products = soup.find_all('div', class_='product') print(f"从{url}获取到{len(products)}个商品") with open(filename, "a", encoding='utf-8') as f: for product in products: # 提取商品标题 title_elem = product.find('h4', class_='product__title') product_title = title_elem.text.strip() if title_elem else "无标题" # 提取当前价格 price_elem = product.find('div', class_='price') current_price = price_elem['data-price'] if price_elem else "无价格" # 提取原价(处理缺失情况) deleted_price_elem = product.find('div', class_='deleted__price') deleted_price = deleted_price_elem.text.strip() if deleted_price_elem else "无原价" # 写入数据 f.write(f"{datetime.now().strftime('%Y-%m-%d %H:%M:%S')}, {product_title}, {current_price}, {deleted_price}\n") print(f"{url}的商品信息已追加到文件!") print(f"所有数据已保存到桌面文件:{filename}")
关键修改说明
- 添加
User-Agent请求头,避免被网站反爬拦截 - 对每个元素的提取增加判断,防止因元素缺失导致代码中断
- 将文件名生成移到循环外,所有页面数据写入同一个文件,方便查看
- 增加请求状态码检查,便于排查请求失败问题
- 写入文件时指定
utf-8编码,避免中文/特殊字符乱码
内容的提问来源于stack exchange,提问作者Koutsa Koutsa
相关产品推荐
相关产品推荐

