使用Python Beautiful Soup爬取Mercado Livre图片时Base64格式问题求助
解决Mercado Livre爬取时图片返回Base64占位图的问题
问题描述
爬取Mercado Livre商品数据时,获取到的图片是Base64格式的占位GIF:
"image": "data:image/gif;base64,R0lGODlhAQABAIAAAAAAAP///yH5BAEAAAAALAAAAAABAAEAAAIBRAA7",
但实际需要的是类似https://http2.mlstatic.com/D_NQ_NP_609104-MLA50695427900_072022-V.webp的真实图片URL。
你的爬取代码片段:
def search_mercadolivre_by_category(category): url = f"https://lista.mercadolivre.com.br/{category}" response = requests.get(url) soup = BeautifulSoup(response.content, 'html.parser') products = soup.find_all("li", {"class": "ui-search-layout__item"}) results = [] for product in products: title = product.find("h2", {"class": "ui-search-item__title"}).text.strip() price = product.find("span", {"class": "price-tag-fraction"}).text.strip() link = product.find("a", {"class": "ui-search-link"})['href'] image = product.find("img")['src'] results.append({ "title": title, "price": price, "link": link, "image": image, "category": category, "website": "Mercado Livre", "keyword": "" }) return results
解决方案
原因分析
Mercado Livre采用图片懒加载机制:页面初始加载时,img标签的src属性指向极小的Base64占位GIF,真实图片URL存储在data-src或data-srcset这类属性中。
修改后的代码
优先读取data-src属性,当该属性不存在时再 fallback 到src:
def search_mercadolivre_by_category(category): url = f"https://lista.mercadolivre.com.br/{category}" response = requests.get(url) soup = BeautifulSoup(response.content, 'html.parser') products = soup.find_all("li", {"class": "ui-search-layout__item"}) results = [] for product in products: title = product.find("h2", {"class": "ui-search-item__title"}).text.strip() price = product.find("span", {"class": "price-tag-fraction"}).text.strip() link = product.find("a", {"class": "ui-search-link"})['href'] # 优先获取真实图片地址的data-src属性,无则用src img_tag = product.find("img") image = img_tag.get('data-src') or img_tag.get('src') results.append({ "title": title, "price": price, "link": link, "image": image, "category": category, "website": "Mercado Livre", "keyword": "" }) return results
补充处理逻辑
如果data-src也无法获取,可检查data-srcset属性(通常包含多分辨率图片地址,取第一个即可):
img_tag = product.find("img") srcset = img_tag.get('data-srcset') if srcset: # 拆分srcset字符串,提取第一个图片URL image = srcset.split(',')[0].split(' ')[0] else: image = img_tag.get('data-src') or img_tag.get('src')
注意:网站HTML结构可能随时间调整,若后续再次出现问题,需重新查看img标签的属性分布。
内容的提问来源于stack exchange,提问作者Laura
相关产品推荐
相关产品推荐

