Django应用:如何从Amazon商品页面链接获取对应图片?
Hey there! Great question—grabbing Amazon product images directly from their product URLs in a Django app is totally doable, and I’ve got a couple of solid approaches to share with you.
Approach 1: Extract ASIN & Construct Image URL (No Web Scraping Needed)
Amazon product URLs always include an ASIN (Amazon Standard Identification Number)—a 10-character alphanumeric code unique to each product. For your example URL, the ASIN is B07PQ7CRBH (you’ll spot it right after dp/ in the URL).
Most Amazon main product images follow a standardized URL pattern we can build using this ASIN. Here’s how to implement this in Django:
Step 1: Extract the ASIN from the URL
Use regex to pull the ASIN—this handles most common Amazon URL formats:
import re def extract_amazon_asin(product_url): # Match ASIN after "dp/" (most common format) asin_match = re.search(r'dp/([A-Z0-9]{10})', product_url) if asin_match: return asin_match.group(1) # Fallback for other URL structures (e.g., short links) asin_match = re.search(r'/([A-Z0-9]{10})/', product_url) return asin_match.group(1) if asin_match else None
Step 2: Build the image URL
Once you have the ASIN, construct the high-res image URL using Amazon’s standard format. Adjust the size parameter (like _SL1500) to get different dimensions:
def build_amazon_image_url(asin): if not asin: return None # _SL1500 sets width to 1500px; replace with _SL500, _SL300, etc. as needed return f"https://images-na.ssl-images-amazon.com/images/I/{asin}._SL1500.jpg"
⚠️ Note: This works for ~90% of products, but a small number use custom filenames instead of the ASIN. For those cases, use the scraping approach below.
Approach 2: Web Scraping (More Reliable, But Watch For Anti-Crawling)
If the ASIN method fails, scrape the product page directly to pull the image URL. Use requests to fetch the page and BeautifulSoup to parse the HTML.
First, install dependencies
Run this in your terminal:
pip install requests beautifulsoup4
Step 1: Write the scraping function
Add a browser-like user agent to avoid getting blocked by Amazon’s anti-crawling systems:
import requests from bs4 import BeautifulSoup def scrape_amazon_main_image(product_url): # Mimic a real browser request headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } try: response = requests.get(product_url, headers=headers) response.raise_for_status() # Raise error for HTTP issues (e.g., 404, 503) soup = BeautifulSoup(response.text, 'html.parser') # Target the main product image (Amazon uses consistent IDs/classes) main_img = soup.find('img', id='landingImage') or soup.find('img', class_='imgTagWrapper imgTagWrapper-noScript') if main_img and 'src' in main_img.attrs: # Clean up the URL to get a high-res version base_img_url = main_img['src'].split('._')[0] return f"{base_img_url}._SL1500.jpg" return None except Exception as e: print(f"Scraping failed: {str(e)}") return None
Step 3: Combine both methods in Django
For best results, try the ASIN method first, then fall back to scraping if it fails. Here’s an example in a Django view:
from django.http import JsonResponse from .utils.amazon_tools import extract_amazon_asin, build_amazon_image_url, scrape_amazon_main_image def get_product_image(request): # Replace this with your actual product URL from the JSON product_url = "https://www.amazon.in/Samsung-Galaxy-Storage-Additional-Exchange/dp/B07PQ7CRBH/ref=sr_1_11?keywords=phone&qid=1563166792&s=electronics&smid=A14CZOWI0VEHLG&sr=1-11" asin = extract_amazon_asin(product_url) image_url = build_amazon_image_url(asin) # Fall back to scraping if ASIN method didn't work if not image_url: image_url = scrape_amazon_main_image(product_url) return JsonResponse({'product_image_url': image_url})
Important Scraping Notes:
- Amazon’s anti-crawling systems may block frequent requests. Add delays between requests or use proxies if processing large numbers of URLs.
- Amazon occasionally updates page structures—you may need to adjust the BeautifulSoup selectors if scraping stops working.
内容的提问来源于stack exchange,提问作者bishwas

