You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

本地Python代码调用Databricks托管LLM API所需PAT最小权限范围咨询

Alright, let's break down the minimal PAT permissions you need for your specific use case—calling a Databricks-hosted LLM via the API for text summarization, without overgranting unnecessary access.

Minimal PAT Permissions for Your LLM Summarization Workflow

Your core task is invoking a pre-deployed Databricks Model Serving endpoint (since you're using the Databricks SDK to interact with the hosted LLM). Here's the tight, minimal set of permissions you require:

  • model-serving/endpoints:invoke: This is non-negotiable—it grants your PAT permission to send inference requests to your target LLM endpoint. Without this, you'll hit a 403 Forbidden error every time you try to run summarization.
  • Optional (for validation only): model-serving/endpoints:get – if you want to check if the endpoint is active or fetch its metadata before making inference calls, but this isn't required for the actual summarization workflow.

Permissions You Can Safely Ignore

Let's eliminate the permissions that have zero relevance to your task:

  • cluster/command-execution: You're not spinning up clusters or running notebook commands—you're only calling a pre-deployed model endpoint, so these permissions are totally unnecessary.
  • Broad mlflow permissions: Unless you're registering new models or modifying the model registry, you don't need MLflow-specific scopes. The model-serving/endpoints:invoke permission already covers access to the registered model behind your endpoint.
  • Any admin/wide-scoped permissions: You already know these pose security risks, so we can rule them out entirely.

Testing Your Configured Permissions

To confirm your PAT has the right access:

  1. Create a new PAT with only the model-serving/endpoints:invoke permission (add get if you want validation capabilities).
  2. Run a quick test with the Databricks SDK to verify:
    from databricks.sdk import WorkspaceClient
    import os
    from dotenv import load_dotenv
    
    load_dotenv()
    w = WorkspaceClient(host=os.getenv("DATABRICKS_HOST"), token=os.getenv("DATABRICKS_TOKEN"))
    
    # Test invoking your LLM endpoint
    response = w.serving_endpoints.invoke_endpoint(
        endpoint_name="your-llm-endpoint-name",
        request={"inputs": ["Sample text to summarize for testing purposes"]}
    )
    print(response)
    
  3. If you get a successful inference response, your permissions are correctly set. If you hit a 403, double-check that the PAT is assigned the model-serving/endpoints:invoke permission and that you're targeting the correct endpoint name.

Quick Note on OpenAI SDK Compatibility

If you ever switch back to using the OpenAI SDK with Databricks' OpenAI-compatible API, the same model-serving/endpoints:invoke permission applies—under the hood, it's still making calls to the Model Serving endpoint.

内容的提问来源于stack exchange,提问作者Mickle-The-Pickle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 10:49:12