网站产品图片爬取不全及与CSV产品关联错误的技术求助(适配Magento2导入)
Fixing Your Product Scraping & Image Association Issues
Let's break down your two core problems and fix the code to get you ready for Magento 2 import today:
Problem Diagnosis
1. Wrong Image-Product Association
Your imagesList is a global variable—you never reset it between products. That means every new product you scrape will inherit all images from every previous product, creating duplicate product rows with unrelated images. Magento 2 expects one row per product, with multiple images separated by commas in the images column, so this is a critical issue.
2. Incomplete Image Downloads
- You're using BeautifulSoup
Tagobjects as keys in yourddictionary, which doesn't work. Eachimageis a unique Tag instance, even if the source URL is identical—so your "duplicate check" does nothing, and you might be skipping valid images by accident. - No error handling for network requests: if a single image download fails (timeout, 404, etc.), your entire script crashes, leaving you with incomplete images.
- Switching directories with
os.chdir(folder)is risky—if the script crashes mid-run, you'll be stuck in theimages_testfolder, which can cause issues with subsequent file operations.
Fixed Code
import requests from bs4 import BeautifulSoup import pandas as pd import os from requests.exceptions import RequestException import re baseurl = 'https://www.k-beauty.ro/magazin/' productLinks = [] # Gather all product links first for page_num in range(1, 8): try: r = requests.get(f'{baseurl}page/{page_num}', timeout=10) r.raise_for_status() # Catch HTTP errors like 404/500 soup = BeautifulSoup(r.content, 'lxml') product_containers = soup.find_all('div', class_='product-element-top') for container in product_containers: for link in container.find_all('a', href=True, class_='product-image-link'): productLinks.append(link['href']) except RequestException as e: print(f"Failed to fetch page {page_num}: {str(e)}") continue ProductItemsList = [] image_folder = 'images_test' # Ensure image folder exists (no need to switch directories) os.makedirs(image_folder, exist_ok=True) # Track downloaded image URLs to avoid duplicates downloaded_image_urls = set() # Process each product page for product_url in productLinks: try: r = requests.get(product_url, timeout=10) r.raise_for_status() soup = BeautifulSoup(r.content, 'lxml') # Extract core product data title = soup.find('h1', class_='product_title').text.strip() raw_price = soup.find('p', class_='price').text.strip() short_desc = soup.find('div', class_='woocommerce-product-details__short-description').text.strip() sku = soup.find('span', class_='sku').text.strip() categories = soup.find('span', class_='posted_in').text.strip() full_desc = soup.find('div', class_='wc-tab-inner').text.strip() brand = soup.find('div', id='tab-pwb_tab-content').text.strip() # Extract and download images for THIS product only (reset list per product) product_image_names = [] product_images = soup.find_all('img', class_='attachment-woocommerce_thumbnail size-woocommerce_thumbnail') for img_idx, img_tag in enumerate(product_images, 1): img_src = img_tag.get('src') # Skip if no URL or already downloaded if not img_src or img_src in downloaded_image_urls: continue # Use SKU + index for image names (unique and product-linked) img_name = f"{sku}_{img_idx}.jpg" img_save_path = os.path.join(image_folder, img_name) try: img_response = requests.get(img_src, timeout=10) img_response.raise_for_status() with open(img_save_path, 'wb') as f: f.write(img_response.content) product_image_names.append(img_name) downloaded_image_urls.add(img_src) print(f"Downloaded: {img_name}") except RequestException as e: print(f"Failed to download {img_src}: {str(e)}") continue # Format product data for Magento 2 (one row per product, images comma-separated) # Fix price extraction to avoid truncation price_match = re.search(r'\d+\.?\d*', raw_price) clean_price = price_match.group() if price_match else '' product = { 'sku': sku, 'categories': categories, 'name': title, 'description': full_desc, 'short_description': short_desc, 'price': clean_price, 'images': ','.join(product_image_names) # Magento 2 compatible format } ProductItemsList.append(product) print(f"Processed product: {title}") except RequestException as e: print(f"Failed to process {product_url}: {str(e)}") continue # Save final CSV output_csv = './K_Beauty_test.csv' df = pd.DataFrame(ProductItemsList) print(df.head()) df.to_csv(output_csv, index=False, encoding='utf-8-sig') # UTF-8 with BOM for Excel compatibility print(f"CSV saved to: {output_csv}")
Key Fixes Explained
Correct Image-Product Mapping
- Moved
product_image_namesinside the product loop—this resets the list for every product, ensuring only that product's images are linked to it. - Stored images as a comma-separated string in the
imagescolumn, which is the standard format Magento 2 expects for multiple images per product.
- Moved
Complete & Reliable Image Downloads
- Replaced the broken Tag-based duplicate check with a set of downloaded image URLs—this properly avoids re-downloading the same image across products.
- Added timeout and error handling for all network requests: if one image or page fails, the script keeps running instead of crashing.
- Used
os.makedirswithexist_ok=Trueto ensure the image folder exists without switching directories, eliminating directory-related bugs. - Renamed images using the product SKU + index—this makes it easy to trace images back to products if you need to debug later.
Improved Price Extraction
- Swapped the fragile
price[:5]truncation with a regex that extracts the numeric price value, so you don't end up with partial prices like "199.9" instead of "199.99".
- Swapped the fragile
Quick Extra Tips
- If you notice missing images, check if the page uses other image classes (like
attachment-shop_singlefor main product images). Update thefind_allcall to include them:product_images = soup.find_all('img', {'class': ['attachment-woocommerce_thumbnail', 'attachment-shop_single']}) - For large-scale scraping, consider adding a small delay between requests (e.g.,
time.sleep(1)) to avoid overwhelming the server.
内容的提问来源于stack exchange,提问作者Dan Croitoriu
相关产品推荐
相关产品推荐

