You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Beautiful Soup爬取Mercado Livre图片时Base64格式问题求助

解决Mercado Livre爬取时图片返回Base64占位图的问题

问题描述

爬取Mercado Livre商品数据时,获取到的图片是Base64格式的占位GIF:

"image": "data:image/gif;base64,R0lGODlhAQABAIAAAAAAAP///yH5BAEAAAAALAAAAAABAAEAAAIBRAA7",

但实际需要的是类似https://http2.mlstatic.com/D_NQ_NP_609104-MLA50695427900_072022-V.webp的真实图片URL。

你的爬取代码片段:

def search_mercadolivre_by_category(category):
    url = f"https://lista.mercadolivre.com.br/{category}"
    response = requests.get(url)
    soup = BeautifulSoup(response.content, 'html.parser')
    products = soup.find_all("li", {"class": "ui-search-layout__item"})
    results = []
    for product in products:
        title = product.find("h2", {"class": "ui-search-item__title"}).text.strip()
        price = product.find("span", {"class": "price-tag-fraction"}).text.strip()
        link = product.find("a", {"class": "ui-search-link"})['href']
        image = product.find("img")['src']
        results.append({
            "title": title,
            "price": price,
            "link": link,
            "image": image,
            "category": category,
            "website": "Mercado Livre",
            "keyword": ""
        })
    return results

解决方案

原因分析

Mercado Livre采用图片懒加载机制:页面初始加载时,img标签的src属性指向极小的Base64占位GIF,真实图片URL存储在data-src或data-srcset这类属性中。

修改后的代码

优先读取data-src属性,当该属性不存在时再 fallback 到src:

def search_mercadolivre_by_category(category):
    url = f"https://lista.mercadolivre.com.br/{category}"
    response = requests.get(url)
    soup = BeautifulSoup(response.content, 'html.parser')
    products = soup.find_all("li", {"class": "ui-search-layout__item"})
    results = []
    for product in products:
        title = product.find("h2", {"class": "ui-search-item__title"}).text.strip()
        price = product.find("span", {"class": "price-tag-fraction"}).text.strip()
        link = product.find("a", {"class": "ui-search-link"})['href']
        # 优先获取真实图片地址的data-src属性,无则用src
        img_tag = product.find("img")
        image = img_tag.get('data-src') or img_tag.get('src')
        results.append({
            "title": title,
            "price": price,
            "link": link,
            "image": image,
            "category": category,
            "website": "Mercado Livre",
            "keyword": ""
        })
    return results

补充处理逻辑

如果data-src也无法获取,可检查data-srcset属性(通常包含多分辨率图片地址,取第一个即可):

img_tag = product.find("img")
srcset = img_tag.get('data-srcset')
if srcset:
    # 拆分srcset字符串,提取第一个图片URL
    image = srcset.split(',')[0].split(' ')[0]
else:
    image = img_tag.get('data-src') or img_tag.get('src')

注意:网站HTML结构可能随时间调整,若后续再次出现问题,需重新查看img标签的属性分布。

内容的提问来源于stack exchange,提问作者Laura

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 19:52:46