You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

电商项目4k+产品图片批量搜索及链接Excel存储Python自动化脚本需求

Batch Product Image Search & Excel Export with Python

Got it, I’ve helped several e-commerce devs solve exactly this problem—scaling up image searches for thousands of products without the manual grind. Here’s a practical, step-by-step solution you can adapt right away.

Step 1: Install Required Libraries

First, let’s get the dependencies set up. Open your terminal and run:

pip install requests beautifulsoup4 pandas python-dotenv
  • requests: Sends HTTP requests to image search engines
  • beautifulsoup4: Parses HTML to extract image links
  • pandas: Handles Excel file reading/writing
  • python-dotenv: (Optional) Stores configuration like user agents to avoid hardcoding

Step 2: Core Script Implementation

We’ll use Bing Image Search here (it’s less aggressive with anti-scraping than Google for small-to-medium scale tasks). The script will:

  • Take a list of product names
  • Search each name for relevant images
  • Collect top image links
  • Save everything to an Excel file with product names mapped to their links

Here’s the code with comments to explain each part:

import requests
from bs4 import BeautifulSoup
import pandas as pd
import time
from random import randint
import json

# Configure your user agent (replace with your browser's UA to avoid blocks)
USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
HEADERS = {"User-Agent": USER_AGENT}

def search_product_images(product_name, num_images=3):
    """Search Bing Images for a product and return top image links"""
    search_url = f"https://www.bing.com/images/search?q={product_name.replace(' ', '+')}&form=HDRSC2"
    image_links = []
    
    try:
        # Add random delay to avoid rate limiting
        time.sleep(randint(2, 5))
        response = requests.get(search_url, headers=HEADERS)
        response.raise_for_status()
        
        soup = BeautifulSoup(response.text, "html.parser")
        # Find all image result elements
        image_elements = soup.find_all("a", class_="iusc")
        
        for elem in image_elements[:num_images]:
            # Extract the image URL from the element's data attributes
            img_data = elem.get("m")
            if img_data:
                img_info = json.loads(img_data)
                img_link = img_info.get("murl")
                if img_link:
                    image_links.append(img_link)
                    
    except requests.exceptions.RequestException as e:
        print(f"Error searching for {product_name}: {e}")
    
    return image_links

def main():
    # Replace this with your list of product names (or read from Excel/CSV)
    # Example: Read from existing Excel
    # df = pd.read_excel("your_product_list.xlsx")
    # product_names = df["Product Name"].tolist()
    
    product_names = [
        "Wireless Bluetooth Headphones",
        "Stainless Steel Water Bottle",
        "Portable Charger 20000mAh",
        # Add your 4000+ products here
    ]
    
    # Collect image links for each product
    results = []
    total_products = len(product_names)
    for idx, product in enumerate(product_names, 1):
        print(f"Processing product {idx}/{total_products}: {product}")
        links = search_product_images(product)
        # Store product name and its image links (join links with commas for Excel)
        results.append({
            "Product Name": product,
            "Image Links": ", ".join(links) if links else "No links found"
        })
    
    # Save results to Excel
    df = pd.DataFrame(results)
    df.to_excel("product_image_links.xlsx", index=False)
    print("Success! Results saved to product_image_links.xlsx")

if __name__ == "__main__":
    main()

Step 3: Adaptations for Your Use Case

  • Reading Product Names from Excel: If your product list is already in an Excel file, replace the hardcoded product_names list with:
    df = pd.read_excel("your_product_list.xlsx")
    product_names = df["Product Name"].tolist()
    
  • Anti-Scraping Tips:
    • Use a rotating list of user agents if you hit blocks (you can find free UA lists online)
    • Increase the delay between requests (adjust randint(2,5) to higher values like randint(3,6))
    • Split your product list into batches and run the script over multiple sessions to avoid rate limits
  • Higher-Quality Images: To filter for high-res images, add a check for image dimensions in the search_product_images function (you can fetch the image header to get size info before adding it to the list)

Important Notes

  • Always check the search engine's robots.txt file (e.g., https://www.bing.com/robots.txt) to ensure you’re allowed to scrape their image results
  • Some image links might be temporary or lead to low-quality assets—you can add validation steps to skip broken links
  • For 4k+ products, consider running the script overnight to avoid interruptions

内容的提问来源于stack exchange,提问作者Raakesh Vanaraj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 01:27:28