Cloudinary图片转换异常:URL含无效标签及Telegram适配求助
问题描述
我开发了一款对接Telegram的亚马逊爬虫程序,通过Selenium抓取商品页面后,用requests获取商品信息并推送至Telegram频道。为优化帖子视觉效果,计划用Cloudinary修改亚马逊默认商品图片,但遇到两个问题:
- 代码生成的image_url被包裹在
<img src>标签中,而非纯HTTPS URL,导致Telegram无法识别; - 图片处理效果与预期不符:预期左侧显示从亚马逊获取的商品图,但实际生成的图片效果偏差。
问题修正方案
1. 解决<img src>标签包裹URL的问题
Cloudinary的image()方法默认返回完整HTML img标签,要获取纯URL,需改用cloudinary_url()方法,直接生成Telegram可识别的HTTPS链接。
2. 修正图片处理效果偏差问题
当前代码存在核心逻辑错误:用固定资源ID作为图层覆盖商品图,且文本使用占位值,同时图层顺序混乱。修正步骤:
- 直接将亚马逊商品图作为远程图层叠加,而非固定资源;
- 把动态获取的价格、折扣替换到文本图层中,并对特殊字符做URL编码;
- 重新调整变换顺序:先设置画布,再放商品图,最后叠加文本元素。
修正后的完整代码
import cloudinary from cloudinary.utils import cloudinary_url import requests from lxml import html from decimal import Decimal, InvalidOperation import locale # 初始化Cloudinary配置 cloudinary.config( cloud_name = "my_cloud_name", api_key = "my_api_key", api_secret = "my_secret_api" ) def parse_decimal(value, locale='it'): """解析本地化价格字符串为Decimal""" try: locale.setlocale(locale.LC_NUMERIC, locale) return Decimal(locale.atof(value.replace(',', '.'))) except (InvalidOperation, locale.Error): return Decimal(0) def url_from_id(product_id): """生成亚马逊商品页面URL(根据实际域名调整)""" return f"https://www.amazon.it/dp/{product_id}" def get_product_info(product_id): headers = {'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:12.0) Gecko/20100101 Firefox/12.0'} params = {'th': '1', 'psc': '1'} product_page = requests.get(url_from_id(product_id), headers=headers, params=params) product_page_content = html.fromstring(product_page.content) try: # 提取商品基本信息 title = product_page_content.xpath('//span[@id="productTitle"]/text()')[0].strip() old_price = product_page_content.xpath('//span[@data-a-strike="true"]//span[@aria-hidden="true"]/text()')[0].strip() new_price_elem = product_page_content.xpath('//span[contains(translate(@class, "PRICETOPAY", "pricetopay"), "pricetopay")]//span[@class="a-offscreen"]/text()') new_price = new_price_elem[0].strip() if new_price_elem else '' if not new_price or new_price == ' ': new_price_whole = product_page_content.xpath('//span[contains(translate(@class, "PRICETOPAY", "pricetopay"), "pricetopay")]//span[@aria-hidden="true"]//span[@class="a-price-whole"]/text()')[0].strip() new_price_decimal = product_page_content.xpath('//span[contains(translate(@class, "PRICETOPAY", "pricetopay"), "pricetopay")]//span[@aria-hidden="true"]//span[@class="a-price-fraction"]/text()')[0].strip() new_price = f"{new_price_whole},{new_price_decimal}€" # 计算折扣率 old_price_num = parse_decimal(old_price.strip('€')) new_price_num = parse_decimal(new_price.strip('€')) discount_rate = f"-{round(100 - (new_price_num / old_price_num) * 100)}%" if old_price_num > 0 else "0%" # 提取商品图片链接 image = product_page_content.xpath('//img[@id="landingImage"]/@src') if not image: image = product_page_content.xpath('//div[contains(@class, "a-dynamic-image-container")]//img/@src') # 清理图片链接,获取高清原图 image_link = image[0].split("._")[0] + ".jpg" # 构建Cloudinary图片变换参数 transformations = [ # 设置画布尺寸(1000x600) {'width': 1000, 'height': 600, 'crop': 'scale', 'background': '#ffffff'}, # 叠加亚马逊商品图到左侧(500x500,左边距100) {'overlay': {'url': image_link}, 'width': 500, 'height': 500, 'crop': 'limit'}, {'flags': 'layer_apply', 'gravity': 'west', 'x': 100, 'y': 50}, # 叠加新价格到右侧上方 {'color': "#000000", 'overlay': { 'font_family': "roboto", 'font_size': 100, 'font_weight': "bold", 'text': new_price.replace('€', '%E2%82%AC').replace(',', '%2C') }}, {'flags': 'layer_apply', 'gravity': 'north_east', 'x': 50, 'y': 50}, # 叠加旧价格到右侧下方(灰色划线) {'color': "#333333", 'overlay': { 'font_family': "roboto", 'font_size': 90, 'font_weight': "bold", 'text': old_price.replace('€', '%E2%82%AC').replace(',', '%2C') }}, {'flags': 'layer_apply', 'gravity': 'south_east', 'x': 50, 'y': 100}, {'effect': 'underline'}, # 叠加折扣率到右上角红色斜体 {'color': "#FF0000", 'overlay': { 'font_family': "roboto", 'font_size': 100, 'font_weight': "bold", 'font_style': "italic", 'text': discount_rate.replace('%', '%25') }}, {'flags': 'layer_apply', 'gravity': 'north_east', 'x': 50, 'y': 150} ] # 生成纯HTTPS图片URL image_url, _ = cloudinary_url( 'placeholder.jpg', transformation=transformations, secure=True ) print(image_url) return { "product_id": product_id, "title": title, "old_price": old_price, "new_price": new_price, "discount_rate": discount_rate, "image_link": image_url } except Exception as e: print(f"\nError for product id:\n\n{product_id}\n\nbecause:\n\n{str(e)}\n Probably strange formatting of webpage.\n") return None
关键修改说明
- 替换
CloudinaryImage().image()为cloudinary_url(),直接输出纯HTTPS链接; - 用亚马逊商品图的远程URL作为图层,替代固定资源ID;
- 动态替换文本图层内容,对
€、,、%做URL编码; - 调整图层顺序和位置参数,保证商品图在左侧、文本在对应位置;
- 优化价格解析逻辑,增加异常处理;
- 补充商品URL生成函数(需根据实际亚马逊域名调整)。
内容的提问来源于stack exchange,提问作者user25503147
相关产品推荐
相关产品推荐

