使用BeautifulSoup提取谷歌商品搜索页原图及爬取不全的问题
_image_src Attribute Values for Google Shopping Product Images with BeautifulSoup Got it, let's break this down. When scraping Google Shopping, the original product image URLs are tucked away in a custom _image_src attribute (instead of the usual src tag, which often points to a resized thumbnail). Here's exactly how to pull those full-size image links using BeautifulSoup:
Step 1: Fetch the Page with Proper Request Headers
First, you need to mimic a real browser to avoid getting blocked by Google's anti-scraping measures. Use the requests library with a valid User-Agent header:
import requests from bs4 import BeautifulSoup # Replace with your browser's User-Agent (find it via "what's my user agent" online) headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # Your target Google Shopping URL url = "https://www.google.com/search?biw=1046&bih=720&tbm=shop&ei=sznVWvq5OcbrzgKPgKLoDA&q=red+dress&oq=red+dress&gs_l=psy-ab.3..0l10.1256.2298.0.2485.9.7.0.0.0.0.238.408.0j1j1.2.0....0...1c.1.64.psy-ab..7.2.407....0.WHO8-4Nhfj0" # Get the page content response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser")
Step 2: Locate Elements with the _image_src Attribute
Google's product images are wrapped in <img> tags that include the _image_src attribute. Use BeautifulSoup's find_all method to target these elements specifically:
# Find all img tags that have the _image_src attribute product_img_tags = soup.find_all("img", attrs={"_image_src": True}) # Extract the attribute values (the original image URLs) original_image_urls = [img["_image_src"] for img in product_img_tags] # Print out the results to verify for idx, img_url in enumerate(original_image_urls, 1): print(f"Product Image {idx}: {img_url}")
Bonus: Decode Encoded URLs (If Needed)
Occasionally, the _image_src value might be URL-encoded. If you notice weird characters like %2F instead of /, use urllib.parse.unquote to decode them:
from urllib.parse import unquote decoded_image_urls = [unquote(url) for url in original_image_urls]
A Quick Heads-Up
Google frequently updates its page structure, so if this selector stops working later, re-inspect the page elements to check if the attribute name or parent tags have changed.
内容的提问来源于stack exchange,提问作者Bob

