You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup/Selenium下载网页非显性图片?

欧泊图片下载与数据整理问题解决方案

问题背景

  • 正在做学习项目,需整理欧泊的名称、价格列表,并从目标网站下载对应图片
  • 已成功获取名称+价格列表,但图片下载遇阻:图片无.jpg后缀、无src属性,提取的链接无法直接使用(示例链接:/storage/images/image?remote=https%3A%2F%2Fwww.koroit-opal-company.com%2FWebRoot%2FStore15%2FShops%2F80300026%2F6364%2FF352%2FFF36%2F99BC%2F5F60%2F0A0C%2F6D0B%2FF3AC%2FIMG-8248VOLL_5_ECK.JPG&shop=80300026&width={width}&height=2560)
  • 补充:已修改代码成功提取图片列表,但仍无法实现下载

网站图片结构示例

enter image description here

当前代码

import requests
from bs4 import BeautifulSoup

URL = 'https://www.koroit-opal-company.com/en/c/solid-opals'
page = requests.get(URL)
headers = {"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36"}
soup = BeautifulSoup(requests.get(f'https://www.koroit-opal-company.com/en/c/solid-opals', headers=headers).text, "lxml")
# print(soup.prettify())

names = [n.getText(strip=True) for n in soup.select("div div div a h2")]
# print(names)
prices = [n.getText(strip=True) for n in soup.select("div div div h3")]
# print(prices)

for name, price in zip(names, prices):
    print(f"{name} {price}")

opal_list = soup.find('div', attrs = {'class':'content'})    #gives just fragment of website where opal imgs are
imgs = opal_list.find_all('img')
print(imgs)

example = imgs[0]
x = example.attrs['data-src']
print(x)

result = soup.find_all(lambda tag: tag.name == 'img' and
                       tag.get('class') == ['product-item-image'])

print(result)

解决方案

1. 图片链接解析逻辑

你提取的data-src是动态加载的中转链接,其中remote参数就是加密后的真实图片地址,需要通过URL解码获取原始链接:

  • 使用urllib.parse模块解析data-src的查询参数
  • 对remote参数值进行URL解码,得到可直接访问的图片URL

2. 完整修改代码

import requests
from bs4 import BeautifulSoup
from urllib.parse import urlparse, parse_qs, unquote
import os

# 目标页面URL
URL = 'https://www.koroit-opal-company.com/en/c/solid-opals'
headers = {"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36"}

# 获取页面内容
response = requests.get(URL, headers=headers)
soup = BeautifulSoup(response.text, "lxml")

# 提取名称和价格
names = [n.getText(strip=True) for n in soup.select("div div div a h2")]
prices = [n.getText(strip=True) for n in soup.select("div div div h3")]

# 创建图片保存目录
if not os.path.exists('opal_images'):
    os.makedirs('opal_images')

# 遍历下载图片
for idx, (name, price) in enumerate(zip(names, prices)):
    # 获取对应产品的图片元素
    product_img = soup.find_all('img', class_='product-item-image')[idx]
    data_src = product_img.get('data-src')
    
    # 解析真实图片URL
    parsed_link = urlparse(data_src)
    query_params = parse_qs(parsed_link.query)
    real_img_url = unquote(query_params['remote'][0])
    
    # 生成合法文件名(替换非法字符)
    safe_name = name.replace('/', '_').replace(':', '_').replace('\\', '_')
    safe_price = price.replace('€', 'EUR').replace(',', '.')
    filename = f"{safe_name}_{safe_price}.jpg"
    save_path = os.path.join('opal_images', filename)
    
    # 下载并保存图片
    img_response = requests.get(real_img_url, headers=headers)
    if img_response.status_code == 200:
        with open(save_path, 'wb') as f:
            f.write(img_response.content)
        print(f"已保存:{filename}")
    else:
        print(f"下载失败:{real_img_url}")

代码说明

  • 自动创建opal_images文件夹存放图片,避免路径混乱
  • 对文件名中的非法字符进行替换,防止保存出错
  • 一一对应名称、价格与图片,确保数据关联正确
  • 增加下载状态提示,方便排查问题

内容的提问来源于stack exchange,提问作者michalb93

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 20:25:19