You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从JSON数据中提取产品JPG图片URL及Alt属性

爬虫需求:提取JSON中指定格式图片URL及对应alt字段

我是爬虫新手,需要从指定JSON结构里提取所有产品的仅JPG格式图片URL,同时获取对应的alt字段名称。具体路径是:attributes > media_map > (b、c、d等动态键) > src下的lg、xl、xxl规格地址。目前已经用Python代码拿到了media_map,但不知道怎么筛选JPG格式的URL,而且不同产品的media_map里的键数量不一样。

现有代码片段

contents = []
with open('urls.csv','r') as csvf: # Open file in read mode
    urls = csv.reader(csvf)
    for url in urls:
        contents.append(url) # Add each url to list contents
        newlist = []
        for url in contents:
            try:
                page = urlopen(url[0]).read()
                soup = BeautifulSoup(page, 'html.parser')
                scripts = soup.find_all('script')[7].text.strip()[24:]
                data = json.loads(scripts)
                link = data['product']['response']['product']['data']['attributes']['media_map']

JSON示例片段

"a218": {
                "label": "Shape",
                "field_type": "button_select",
                "value_order": ["v766","v767"],
                "values": {
                  "v766": {"label": "Round","value": "S6CBRO","price": 35},
                  "v767": {"label": "Rectangle","value": "S6CBRE","price": 35,"hypotheticalPrice": 24.5}
                }
              },
            "inventory": {"stock": 0,"sold": 0,"total": 0},
            "optional": {},
            "media_map": {
              "b": {
                "src": {
                  "xs": "https://ctl.s6img.com/society6/img/xVx1vleu7iLcR79ZkRZKqQiSzZE/w_125/artwork/~artwork/s6-0041/a/18613683_5971445",
                  "lg": "https://ctl.s6img.com/society6/img/W-ESMqUtC_oOEUjx-1E_SyIdueI/w_550/artwork/~artwork/s6-0041/a/18613683_5971445",
                  "xl": "https://ctl.s6img.com/society6/img/z90VlaYwd8cxCqbrZ1ttAxINpaY/w_700/artwork/~artwork/s6-0041/a/18613683_5971445",
                  "xxl": null
                },
                "type": "image",
                "alt": "I'M NOT ALWAYS A BITCH (Red) Cutting Board",
                "meta": null
              },
              "c": {
                "src": {
                  "xs": "https://ctl.s6img.com/society6/img/KQJbb4jG0gBHcqQiOCivLUbKMxI/w_125/cutting-board/rectangle/lifestyle/~artwork,fw_1572,fh_2500,fx_93,fy_746,iw_1386,ih_2142/s6-0041/a/18613725_13086827/~~/im-not-always-a-bitch-red-cutting-board.jpg",
                  "lg": "https://ctl.s6img.com/society6/img/ztGrxSpA7FC1LfzM3UldiQkEi7g/w_550/cutting-board/rectangle/lifestyle/~artwork,fw_1572,fh_2500,fx_93,fy_746,iw_1386,ih_2142/s6-0041/a/18613725_13086827/~~/im-not-always-a-bitch-red-cutting-board.jpg",
                  "xl": "https://ctl.s6img.com/society6/img/PHjp9jDic2NGUrpq8k0aaxsYZr4/w_700/cutting-board/rectangle/lifestyle/~artwork,fw_1572,fh_2500,fx_93,fy_746,iw_1386,ih_2142/s6-0041/a/18613725_13086827/~~/im-not-always-a-bitch-red-cutting-board.jpg",
                  "xxl": "https://ctl.s6img.com/society6/img/m-1HhSM5CIGl6DY9ukCVxSmVDIw/w_1500/cutting-board/rectangle/lifestyle/~artwork,fw_1572,fh_2500,fx_93,fy_746,iw_1386,ih_2142/s6-0041/a/18613725_13086827/~~/im-not-always-a-bitch-red-cutting-board.jpg"
                }
              }
            }

解决方案

核心思路

  1. 遍历media_map中的所有动态键(不管是b、c还是其他)
  2. 对每个键对应的内容,提取alt字段(兼容无alt的情况)
  3. 筛选src下lg/xl/xxl规格的URL,只保留以.jpg结尾且不为空的地址

修改后的完整代码

import json
from urllib.request import urlopen
from bs4 import BeautifulSoup
import csv

# 存储最终结果的列表
image_results = []

with open('urls.csv','r') as csvf:
    urls = csv.reader(csvf)
    for url_row in urls:
        target_url = url_row[0]
        try:
            # 获取页面内容并解析JSON数据
            page = urlopen(target_url).read()
            soup = BeautifulSoup(page, 'html.parser')
            script_content = soup.find_all('script')[7].text.strip()[24:]
            data = json.loads(script_content)
            media_map = data['product']['response']['product']['data']['attributes']['media_map']
            
            # 遍历media_map的所有动态键
            for media_key, media_details in media_map.items():
                # 获取alt文本,没有则设为None
                alt_text = media_details.get('alt', None)
                src_list = media_details.get('src', {})
                
                # 筛选指定规格的图片URL
                for size in ['lg', 'xl', 'xxl']:
                    img_url = src_list.get(size)
                    # 检查URL是否存在且为JPG格式
                    if img_url and img_url.endswith('.jpg'):
                        image_results.append({
                            'alt': alt_text,
                            'size': size,
                            'image_url': img_url
                        })
                        
        except Exception as e:
            print(f"处理URL {target_url} 时出错: {str(e)}")

# 打印结果示例
for item in image_results:
    print(f"Alt: {item['alt']}, 规格: {item['size']}, URL: {item['image_url']}")

# 保存结果到CSV文件
with open('image_output.csv', 'w', newline='', encoding='utf-8') as csvfile:
    fieldnames = ['alt', 'size', 'image_url']
    writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
    writer.writeheader()
    writer.writerows(image_results)

关键细节说明

  • 用media_map.items()遍历所有动态键,适配不同产品的键数量差异
  • 用str.endswith('.jpg')精准筛选JPG格式URL
  • 用dict.get()方法避免键不存在时抛出异常(比如部分媒体项无alt、部分规格URL为null)
  • 加入异常处理,单个URL出错不会导致整个程序终止

内容的提问来源于stack exchange,提问作者babar akhter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 14:50:37