You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取谷歌商品搜索页原图及爬取不全的问题

How to Extract _image_src Attribute Values for Google Shopping Product Images with BeautifulSoup

Got it, let's break this down. When scraping Google Shopping, the original product image URLs are tucked away in a custom _image_src attribute (instead of the usual src tag, which often points to a resized thumbnail). Here's exactly how to pull those full-size image links using BeautifulSoup:

Step 1: Fetch the Page with Proper Request Headers

First, you need to mimic a real browser to avoid getting blocked by Google's anti-scraping measures. Use the requests library with a valid User-Agent header:

import requests
from bs4 import BeautifulSoup

# Replace with your browser's User-Agent (find it via "what's my user agent" online)
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# Your target Google Shopping URL
url = "https://www.google.com/search?biw=1046&bih=720&tbm=shop&ei=sznVWvq5OcbrzgKPgKLoDA&q=red+dress&oq=red+dress&gs_l=psy-ab.3..0l10.1256.2298.0.2485.9.7.0.0.0.0.238.408.0j1j1.2.0....0...1c.1.64.psy-ab..7.2.407....0.WHO8-4Nhfj0"

# Get the page content
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

Step 2: Locate Elements with the _image_src Attribute

Google's product images are wrapped in <img> tags that include the _image_src attribute. Use BeautifulSoup's find_all method to target these elements specifically:

# Find all img tags that have the _image_src attribute
product_img_tags = soup.find_all("img", attrs={"_image_src": True})

# Extract the attribute values (the original image URLs)
original_image_urls = [img["_image_src"] for img in product_img_tags]

# Print out the results to verify
for idx, img_url in enumerate(original_image_urls, 1):
    print(f"Product Image {idx}: {img_url}")

Bonus: Decode Encoded URLs (If Needed)

Occasionally, the _image_src value might be URL-encoded. If you notice weird characters like %2F instead of /, use urllib.parse.unquote to decode them:

from urllib.parse import unquote

decoded_image_urls = [unquote(url) for url in original_image_urls]

A Quick Heads-Up

Google frequently updates its page structure, so if this selector stops working later, re-inspect the page elements to check if the attribute name or parent tags have changed.

内容的提问来源于stack exchange,提问作者Bob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:27:39