You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网站产品图片爬取不全及与CSV产品关联错误的技术求助(适配Magento2导入)

Fixing Your Product Scraping & Image Association Issues

Let's break down your two core problems and fix the code to get you ready for Magento 2 import today:

Problem Diagnosis

1. Wrong Image-Product Association

Your imagesList is a global variable—you never reset it between products. That means every new product you scrape will inherit all images from every previous product, creating duplicate product rows with unrelated images. Magento 2 expects one row per product, with multiple images separated by commas in the images column, so this is a critical issue.

2. Incomplete Image Downloads

  • You're using BeautifulSoup Tag objects as keys in your d dictionary, which doesn't work. Each image is a unique Tag instance, even if the source URL is identical—so your "duplicate check" does nothing, and you might be skipping valid images by accident.
  • No error handling for network requests: if a single image download fails (timeout, 404, etc.), your entire script crashes, leaving you with incomplete images.
  • Switching directories with os.chdir(folder) is risky—if the script crashes mid-run, you'll be stuck in the images_test folder, which can cause issues with subsequent file operations.

Fixed Code

import requests
from bs4 import BeautifulSoup
import pandas as pd
import os
from requests.exceptions import RequestException
import re

baseurl = 'https://www.k-beauty.ro/magazin/'
productLinks = []

# Gather all product links first
for page_num in range(1, 8):
    try:
        r = requests.get(f'{baseurl}page/{page_num}', timeout=10)
        r.raise_for_status()  # Catch HTTP errors like 404/500
        soup = BeautifulSoup(r.content, 'lxml')
        product_containers = soup.find_all('div', class_='product-element-top')
        
        for container in product_containers:
            for link in container.find_all('a', href=True, class_='product-image-link'):
                productLinks.append(link['href'])
                
    except RequestException as e:
        print(f"Failed to fetch page {page_num}: {str(e)}")
        continue

ProductItemsList = []
image_folder = 'images_test'
# Ensure image folder exists (no need to switch directories)
os.makedirs(image_folder, exist_ok=True)
# Track downloaded image URLs to avoid duplicates
downloaded_image_urls = set()

# Process each product page
for product_url in productLinks:
    try:
        r = requests.get(product_url, timeout=10)
        r.raise_for_status()
        soup = BeautifulSoup(r.content, 'lxml')
        
        # Extract core product data
        title = soup.find('h1', class_='product_title').text.strip()
        raw_price = soup.find('p', class_='price').text.strip()
        short_desc = soup.find('div', class_='woocommerce-product-details__short-description').text.strip()
        sku = soup.find('span', class_='sku').text.strip()
        categories = soup.find('span', class_='posted_in').text.strip()
        full_desc = soup.find('div', class_='wc-tab-inner').text.strip()
        brand = soup.find('div', id='tab-pwb_tab-content').text.strip()

        # Extract and download images for THIS product only (reset list per product)
        product_image_names = []
        product_images = soup.find_all('img', class_='attachment-woocommerce_thumbnail size-woocommerce_thumbnail')
        
        for img_idx, img_tag in enumerate(product_images, 1):
            img_src = img_tag.get('src')
            # Skip if no URL or already downloaded
            if not img_src or img_src in downloaded_image_urls:
                continue
            
            # Use SKU + index for image names (unique and product-linked)
            img_name = f"{sku}_{img_idx}.jpg"
            img_save_path = os.path.join(image_folder, img_name)
            
            try:
                img_response = requests.get(img_src, timeout=10)
                img_response.raise_for_status()
                with open(img_save_path, 'wb') as f:
                    f.write(img_response.content)
                
                product_image_names.append(img_name)
                downloaded_image_urls.add(img_src)
                print(f"Downloaded: {img_name}")
                
            except RequestException as e:
                print(f"Failed to download {img_src}: {str(e)}")
                continue
        
        # Format product data for Magento 2 (one row per product, images comma-separated)
        # Fix price extraction to avoid truncation
        price_match = re.search(r'\d+\.?\d*', raw_price)
        clean_price = price_match.group() if price_match else ''
        
        product = {
            'sku': sku,
            'categories': categories,
            'name': title,
            'description': full_desc,
            'short_description': short_desc,
            'price': clean_price,
            'images': ','.join(product_image_names)  # Magento 2 compatible format
        }
        ProductItemsList.append(product)
        print(f"Processed product: {title}")
    
    except RequestException as e:
        print(f"Failed to process {product_url}: {str(e)}")
        continue

# Save final CSV
output_csv = './K_Beauty_test.csv'
df = pd.DataFrame(ProductItemsList)
print(df.head())
df.to_csv(output_csv, index=False, encoding='utf-8-sig')  # UTF-8 with BOM for Excel compatibility
print(f"CSV saved to: {output_csv}")

Key Fixes Explained

  1. Correct Image-Product Mapping

    • Moved product_image_names inside the product loop—this resets the list for every product, ensuring only that product's images are linked to it.
    • Stored images as a comma-separated string in the images column, which is the standard format Magento 2 expects for multiple images per product.
  2. Complete & Reliable Image Downloads

    • Replaced the broken Tag-based duplicate check with a set of downloaded image URLs—this properly avoids re-downloading the same image across products.
    • Added timeout and error handling for all network requests: if one image or page fails, the script keeps running instead of crashing.
    • Used os.makedirs with exist_ok=True to ensure the image folder exists without switching directories, eliminating directory-related bugs.
    • Renamed images using the product SKU + index—this makes it easy to trace images back to products if you need to debug later.
  3. Improved Price Extraction

    • Swapped the fragile price[:5] truncation with a regex that extracts the numeric price value, so you don't end up with partial prices like "199.9" instead of "199.99".

Quick Extra Tips

  • If you notice missing images, check if the page uses other image classes (like attachment-shop_single for main product images). Update the find_all call to include them:
    product_images = soup.find_all('img', {'class': ['attachment-woocommerce_thumbnail', 'attachment-shop_single']})
    
  • For large-scale scraping, consider adding a small delay between requests (e.g., time.sleep(1)) to avoid overwhelming the server.

内容的提问来源于stack exchange,提问作者Dan Croitoriu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 20:57:31