本地Python代码调用Databricks托管LLM API所需PAT最小权限范围咨询
Alright, let's break down the minimal PAT permissions you need for your specific use case—calling a Databricks-hosted LLM via the API for text summarization, without overgranting unnecessary access.
Minimal PAT Permissions for Your LLM Summarization Workflow
Your core task is invoking a pre-deployed Databricks Model Serving endpoint (since you're using the Databricks SDK to interact with the hosted LLM). Here's the tight, minimal set of permissions you require:
model-serving/endpoints:invoke: This is non-negotiable—it grants your PAT permission to send inference requests to your target LLM endpoint. Without this, you'll hit a 403 Forbidden error every time you try to run summarization.- Optional (for validation only):
model-serving/endpoints:get– if you want to check if the endpoint is active or fetch its metadata before making inference calls, but this isn't required for the actual summarization workflow.
Permissions You Can Safely Ignore
Let's eliminate the permissions that have zero relevance to your task:
cluster/command-execution: You're not spinning up clusters or running notebook commands—you're only calling a pre-deployed model endpoint, so these permissions are totally unnecessary.- Broad
mlflowpermissions: Unless you're registering new models or modifying the model registry, you don't need MLflow-specific scopes. Themodel-serving/endpoints:invokepermission already covers access to the registered model behind your endpoint. - Any admin/wide-scoped permissions: You already know these pose security risks, so we can rule them out entirely.
Testing Your Configured Permissions
To confirm your PAT has the right access:
- Create a new PAT with only the
model-serving/endpoints:invokepermission (addgetif you want validation capabilities). - Run a quick test with the Databricks SDK to verify:
from databricks.sdk import WorkspaceClient import os from dotenv import load_dotenv load_dotenv() w = WorkspaceClient(host=os.getenv("DATABRICKS_HOST"), token=os.getenv("DATABRICKS_TOKEN")) # Test invoking your LLM endpoint response = w.serving_endpoints.invoke_endpoint( endpoint_name="your-llm-endpoint-name", request={"inputs": ["Sample text to summarize for testing purposes"]} ) print(response) - If you get a successful inference response, your permissions are correctly set. If you hit a 403, double-check that the PAT is assigned the
model-serving/endpoints:invokepermission and that you're targeting the correct endpoint name.
Quick Note on OpenAI SDK Compatibility
If you ever switch back to using the OpenAI SDK with Databricks' OpenAI-compatible API, the same model-serving/endpoints:invoke permission applies—under the hood, it's still making calls to the Model Serving endpoint.
内容的提问来源于stack exchange,提问作者Mickle-The-Pickle

