You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS SageMaker部署HuggingFace Llama模型时自定义inference.py脚本未执行的问题求助

AWS SageMaker部署HuggingFace Llama模型时自定义inference.py脚本未执行的问题求助

大家好,我最近在SageMaker Studio的JupyterLab环境里折腾部署HuggingFace的Llama-3.1-8B-Instruct模型,目标是用来做批量转换推理。我的数据存在同AWS账号的S3桶里,需要做一些自定义处理(比如转JSON格式、包装成LLM需要的prompt等),所以必须用到自定义的inference.py脚本。但不管我怎么试,这个脚本都完全没跑起来,卡了好几天了,求各位大佬帮忙看看!

我的环境与目录结构

  • 运行环境:SageMaker Studio的JupyterLab空间
  • 目录结构:笔记本文件和自定义脚本在同一个EFS子目录下:
    ~
    |---- user-default-efs
          |---- notebook.ipynb
          |---- inference.py
    

部署代码(notebook中)

我用的是HuggingFace官方的LLM镜像,部署代码如下(敏感信息用占位符替换了):

import json
import sagemaker
from sagemaker.huggingface import HuggingFaceModel, get_huggingface_llm_image_uri

hub = {
    "HF_MODEL_ID": "meta-llama/Llama-3.1-8B-Instruct",
    "SM_NUM_GPUS": json.dumps(1),
    "HF_TOKEN": "{api_key_here}"  # 已替换为我的HuggingFace密钥
}

instance_type, instance_count = "{ec2_instance_here}", 1  # 使用的是符合要求的GPU实例

hf_model = HuggingFaceModel(
    image_uri=get_huggingface_llm_image_uri("huggingface", version="3.2.3"),
    env=hub,
    role=sagemaker.get_execution_role(),
    entry_point="inference.py"
)

predictor = hf_model.deploy(
    initial_instance_count=instance_count,
    instance_type=instance_type
)

# 测试用的请求payload
payload = {
    "inputs": "hello world!",
    "parameters": {
        "max_new_tokens": 64,
        "temperature": 0.5,
        "do_sample": True
    }
}

response = predictor.predict(payload)
print(response)

注:目前为了快速测试,先部署了标准endpoint,还没到批量转换的步骤。

自定义inference.py脚本

我写的测试脚本非常简单,完全没做实际的推理逻辑,只是想先确认脚本是否被执行:

def input_fn(input_data, content_type):
    print(f"{input_data=}")
    print(f"{content_type=}")
    return input_data

def predict_fn(processed_data, model):
    print(f"{processed_data=}")
    print(f"{model=}")
    return "This is a test prediction."

def output_fn(prediction, accept):
    print(f"{prediction=}")
    print(f"{accept=}")
    return "This is a test output.", accept

问题现象

  1. 返回结果不符合预期:按照我的设计,调用predict应该返回output_fn里的"This is a test output.",并且完全绕过Llama模型的默认推理,但实际返回的是Llama模型对"hello world!"的正常回复,就好像inference.py根本不存在一样。
  2. 无日志输出:脚本里的所有print语句,既没出现在JupyterLab的笔记本输出中,也没在CloudWatch的日志里找到,这让我严重怀疑脚本根本没被执行。

有没有朋友遇到过类似的问题?能帮我分析下可能的原因吗?或者有什么排查的方向?万分感谢!

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 03:08:52