电商项目4k+产品图片批量搜索及链接Excel存储Python自动化脚本需求
Batch Product Image Search & Excel Export with Python
Got it, I’ve helped several e-commerce devs solve exactly this problem—scaling up image searches for thousands of products without the manual grind. Here’s a practical, step-by-step solution you can adapt right away.
Step 1: Install Required Libraries
First, let’s get the dependencies set up. Open your terminal and run:
pip install requests beautifulsoup4 pandas python-dotenv
requests: Sends HTTP requests to image search enginesbeautifulsoup4: Parses HTML to extract image linkspandas: Handles Excel file reading/writingpython-dotenv: (Optional) Stores configuration like user agents to avoid hardcoding
Step 2: Core Script Implementation
We’ll use Bing Image Search here (it’s less aggressive with anti-scraping than Google for small-to-medium scale tasks). The script will:
- Take a list of product names
- Search each name for relevant images
- Collect top image links
- Save everything to an Excel file with product names mapped to their links
Here’s the code with comments to explain each part:
import requests from bs4 import BeautifulSoup import pandas as pd import time from random import randint import json # Configure your user agent (replace with your browser's UA to avoid blocks) USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" HEADERS = {"User-Agent": USER_AGENT} def search_product_images(product_name, num_images=3): """Search Bing Images for a product and return top image links""" search_url = f"https://www.bing.com/images/search?q={product_name.replace(' ', '+')}&form=HDRSC2" image_links = [] try: # Add random delay to avoid rate limiting time.sleep(randint(2, 5)) response = requests.get(search_url, headers=HEADERS) response.raise_for_status() soup = BeautifulSoup(response.text, "html.parser") # Find all image result elements image_elements = soup.find_all("a", class_="iusc") for elem in image_elements[:num_images]: # Extract the image URL from the element's data attributes img_data = elem.get("m") if img_data: img_info = json.loads(img_data) img_link = img_info.get("murl") if img_link: image_links.append(img_link) except requests.exceptions.RequestException as e: print(f"Error searching for {product_name}: {e}") return image_links def main(): # Replace this with your list of product names (or read from Excel/CSV) # Example: Read from existing Excel # df = pd.read_excel("your_product_list.xlsx") # product_names = df["Product Name"].tolist() product_names = [ "Wireless Bluetooth Headphones", "Stainless Steel Water Bottle", "Portable Charger 20000mAh", # Add your 4000+ products here ] # Collect image links for each product results = [] total_products = len(product_names) for idx, product in enumerate(product_names, 1): print(f"Processing product {idx}/{total_products}: {product}") links = search_product_images(product) # Store product name and its image links (join links with commas for Excel) results.append({ "Product Name": product, "Image Links": ", ".join(links) if links else "No links found" }) # Save results to Excel df = pd.DataFrame(results) df.to_excel("product_image_links.xlsx", index=False) print("Success! Results saved to product_image_links.xlsx") if __name__ == "__main__": main()
Step 3: Adaptations for Your Use Case
- Reading Product Names from Excel: If your product list is already in an Excel file, replace the hardcoded
product_nameslist with:df = pd.read_excel("your_product_list.xlsx") product_names = df["Product Name"].tolist() - Anti-Scraping Tips:
- Use a rotating list of user agents if you hit blocks (you can find free UA lists online)
- Increase the delay between requests (adjust
randint(2,5)to higher values likerandint(3,6)) - Split your product list into batches and run the script over multiple sessions to avoid rate limits
- Higher-Quality Images: To filter for high-res images, add a check for image dimensions in the
search_product_imagesfunction (you can fetch the image header to get size info before adding it to the list)
Important Notes
- Always check the search engine's
robots.txtfile (e.g.,https://www.bing.com/robots.txt) to ensure you’re allowed to scrape their image results - Some image links might be temporary or lead to low-quality assets—you can add validation steps to skip broken links
- For 4k+ products, consider running the script overnight to avoid interruptions
内容的提问来源于stack exchange,提问作者Raakesh Vanaraj
相关产品推荐
相关产品推荐

