You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

部署LaBSE至Google Cloud AI Platform遇响应超2MB限制求解决方案

问题描述

前几日我将LaBSE模型部署到AI Platform,遇到请求返回结果超出2MB大小上限的问题。
我想到了几个优化方案:

  • 让AI Platform返回最小化(非美化格式化)的JSON,去除所有多余空格和换行
  • 让AI Platform以二进制格式返回结果
  • 由于当前返回结果包含约13个输出,将其调整为仅返回1个输出

请问有人知道方案1或方案2的实现方法吗?
我已经在方案3上投入了大量精力,确认该方案可行,比如可以在上传模型前修改网络结构,以下是我目前的尝试:

VERSION = 'v1'
MODEL = 'labse_2_b'
MODEL_DIR = BUCKET + '/' + MODEL

# Download the model
! wget 'https://tfhub.dev/google/LaBSE/2?tf-hub-format=compressed' \
  --output-document='{MODEL}.tar.gz'
! mkdir {MODEL}
! tar -xzvf '{MODEL}.tar.gz' -C {MODEL}

# Attempts to load the model, edit it, and save it:
model.save(export_path, save_format='tf') # ValueError: Model <keras.engine.sequential.Sequential object at 0x7f87e833c650> 
# cannot be saved because the input shapes have not been set. 
# Usually, input shapes are automatically determined from calling 
# `.fit()` or `.predict()`. 
# To manually set the shapes, call `model.build(input_shape)`.
model.build(input_shape=(None,)) # cannot find a proper shape

# create a AI Plateform model version:
! gsutil -m cp -r '{MODEL}' {MODEL_DIR}  # upload model to Google Cloud Storage
! gcloud ai-platform versions create $VERSION \
  --model {MODEL} \
  --origin {MODEL_DIR} \
  --runtime-version=2.1 \
  --framework='tensorflow' \
  --python-version=3.7 \
  --region="{REGION}"

恳请大家帮忙解决该问题,万分感谢!
补充说明:哪怕是仅包含16个分词的短句"I wish you a pleasant flight and a good meal aboard this plane.",对应的分词编码为[101, 146, 34450, 15100, 170, 147508, 48088, 14999, 170, 17072, 66369, 351617, 15272, 69746, 119, 102],也无法正常处理,报错如下:

Response size too large. Received at least 3220082 bytes; max is 2000000.". Details: "Response size too large. Received at least 3220082 bytes; max is 2000000.
解决方案

方案1 最小化JSON返回实现

如果使用AI Platform自定义预测例程,在返回预测结果时使用json.dumps(output, separators=(',', ':'))生成无冗余空格和换行的压缩JSON即可,不要使用默认的美化格式化输出。如果使用平台默认的TensorFlow Serving部署,请求时指定仅返回需要的输出字段,TF Serving默认返回的JSON已经是压缩格式,不会额外增加冗余字符。

方案2 二进制格式返回实现

请求时添加HTTP头Accept: application/x-protobuf,即可直接获取TensorFlow Serving返回的序列化protobuf二进制结果,体积相比同内容JSON小60%以上。如果使用自定义预测代码,也可以将输出张量直接序列化为bytes返回,同时设置响应头Content-Type: application/octet-stream即可。

方案3 现有报错修复

你遇到的模型保存报错是因为没有正确匹配LaBSE的输入结构,LaBSE需要三个输入张量,分别是input_ids、attention_mask、token_type_ids,每个张量的shape为(None, 最大序列长度),可参考以下代码修改:

import tensorflow as tf
import tensorflow_hub as hub

# 加载本地已下载的LaBSE模型
labse_model = hub.load("./labse_2_b")

# 封装为仅返回句向量的单输出模型
class TrimmedLaBSE(tf.keras.Model):
    def __init__(self, labse):
        super().__init__()
        self.labse = labse
    def call(self, inputs):
        # 仅保留需要的句向量输出,丢弃其余12个多余输出
        return self.labse(inputs)["default"]

model = TrimmedLaBSE(labse_model)
# 按实际使用的最大序列长度设置输入shape,示例为最大128个token
model.build([
    tf.TensorShape((None, 128)),
    tf.TensorShape((None, 128)),
    tf.TensorShape((None, 128))
])
# 执行一次测试推理补全模型拓扑
test_input = [tf.zeros((1, 128), dtype=tf.int32) for _ in range(3)]
_ = model(test_input)
# 保存处理后的模型
model.save("./labse_trimmed", save_format="tf")

修改完成后将新生成的labse_trimmed目录上传到GCS创建AI Platform版本即可,单输出的模型单条请求返回体积仅几KB,完全不会触发2MB大小限制。

内容的提问来源于stack exchange,提问作者user2346922

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 15:00:03