You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Django应用:如何从Amazon商品页面链接获取对应图片?

Hey there! Great question—grabbing Amazon product images directly from their product URLs in a Django app is totally doable, and I’ve got a couple of solid approaches to share with you.

Approach 1: Extract ASIN & Construct Image URL (No Web Scraping Needed)

Amazon product URLs always include an ASIN (Amazon Standard Identification Number)—a 10-character alphanumeric code unique to each product. For your example URL, the ASIN is B07PQ7CRBH (you’ll spot it right after dp/ in the URL).

Most Amazon main product images follow a standardized URL pattern we can build using this ASIN. Here’s how to implement this in Django:

Step 1: Extract the ASIN from the URL

Use regex to pull the ASIN—this handles most common Amazon URL formats:

import re

def extract_amazon_asin(product_url):
    # Match ASIN after "dp/" (most common format)
    asin_match = re.search(r'dp/([A-Z0-9]{10})', product_url)
    if asin_match:
        return asin_match.group(1)
    
    # Fallback for other URL structures (e.g., short links)
    asin_match = re.search(r'/([A-Z0-9]{10})/', product_url)
    return asin_match.group(1) if asin_match else None

Step 2: Build the image URL

Once you have the ASIN, construct the high-res image URL using Amazon’s standard format. Adjust the size parameter (like _SL1500) to get different dimensions:

def build_amazon_image_url(asin):
    if not asin:
        return None
    # _SL1500 sets width to 1500px; replace with _SL500, _SL300, etc. as needed
    return f"https://images-na.ssl-images-amazon.com/images/I/{asin}._SL1500.jpg"

⚠️ Note: This works for ~90% of products, but a small number use custom filenames instead of the ASIN. For those cases, use the scraping approach below.

Approach 2: Web Scraping (More Reliable, But Watch For Anti-Crawling)

If the ASIN method fails, scrape the product page directly to pull the image URL. Use requests to fetch the page and BeautifulSoup to parse the HTML.

First, install dependencies

Run this in your terminal:

pip install requests beautifulsoup4

Step 1: Write the scraping function

Add a browser-like user agent to avoid getting blocked by Amazon’s anti-crawling systems:

import requests
from bs4 import BeautifulSoup

def scrape_amazon_main_image(product_url):
    # Mimic a real browser request
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    
    try:
        response = requests.get(product_url, headers=headers)
        response.raise_for_status()  # Raise error for HTTP issues (e.g., 404, 503)
        
        soup = BeautifulSoup(response.text, 'html.parser')
        # Target the main product image (Amazon uses consistent IDs/classes)
        main_img = soup.find('img', id='landingImage') or soup.find('img', class_='imgTagWrapper imgTagWrapper-noScript')
        
        if main_img and 'src' in main_img.attrs:
            # Clean up the URL to get a high-res version
            base_img_url = main_img['src'].split('._')[0]
            return f"{base_img_url}._SL1500.jpg"
        
        return None
    except Exception as e:
        print(f"Scraping failed: {str(e)}")
        return None

Step 3: Combine both methods in Django

For best results, try the ASIN method first, then fall back to scraping if it fails. Here’s an example in a Django view:

from django.http import JsonResponse
from .utils.amazon_tools import extract_amazon_asin, build_amazon_image_url, scrape_amazon_main_image

def get_product_image(request):
    # Replace this with your actual product URL from the JSON
    product_url = "https://www.amazon.in/Samsung-Galaxy-Storage-Additional-Exchange/dp/B07PQ7CRBH/ref=sr_1_11?keywords=phone&qid=1563166792&s=electronics&smid=A14CZOWI0VEHLG&sr=1-11"
    
    asin = extract_amazon_asin(product_url)
    image_url = build_amazon_image_url(asin)
    
    # Fall back to scraping if ASIN method didn't work
    if not image_url:
        image_url = scrape_amazon_main_image(product_url)
    
    return JsonResponse({'product_image_url': image_url})

Important Scraping Notes:

  • Amazon’s anti-crawling systems may block frequent requests. Add delays between requests or use proxies if processing large numbers of URLs.
  • Amazon occasionally updates page structures—you may need to adjust the BeautifulSoup selectors if scraping stops working.

内容的提问来源于stack exchange,提问作者bishwas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 09:52:42