调用带参Databricks Notebook时%run与dbutils.notebook.run的行为差异
问题
使用dbutils.notebook.run()调用Databricks Notebook配置Azure Data Lake Storage Gen2(ADLS)连接后,读取Delta表时报错,但用%run执行相同逻辑却完全正常。
工具Notebook(configure-storage)代码
# Notebook parameters dbutils.widgets.text("storage_account","") dbutils.widgets.text("tenant_id","") dbutils.widgets.text("client_id","") dbutils.widgets.text("client_secret","") # Set storage account and get secrets from Key Vault storage_account = dbutils.widgets.get("storage_account") tenant_id = dbutils.secrets.get(scope="key-vault",key=dbutils.widgets.get("tenant_id")) client_id = dbutils.secrets.get(scope="key-vault",key=dbutils.widgets.get("client_id")) client_secret = dbutils.secrets.get(scope="key-vault",key=dbutils.widgets.get("client_secret")) # Azure Data Lake Storage auth spark.conf.set(f"fs.azure.account.auth.type.{storage_account}.dfs.core.windows.net", "OAuth") spark.conf.set(f"fs.azure.account.oauth.provider.type.{storage_account}.dfs.core.windows.net", "org.apache.hadoop.fs.azurebfs.oauth2.ClientCredsTokenProvider") spark.conf.set(f"fs.azure.account.oauth2.client.id.{storage_account}.dfs.core.windows.net", f"{client_id}") spark.conf.set(f"fs.azure.account.oauth2.client.secret.{storage_account}.dfs.core.windows.net", client_secret) spark.conf.set(f"fs.azure.account.oauth2.client.endpoint.{storage_account}.dfs.core.windows.net", f"https://login.microsoftonline.com/{tenant_id}/oauth2/token")
调用端读取Delta表代码
file_location = "abfss://<storage-container>@<storage-account>.dfs.core.windows.net/<path-to-delta-table>" df = spark.read.format("delta").load(file_location) display(df)
成功的%run调用方式
%run "../util/configure-storage" $storage_account="storage-account-name" $tenant_id="tenant-id-secret-name" $client_id="client-id-secret-name" $client_secret="client-secret-secret-name"
dbutils.notebook.run()调用后的报错信息
Py4JJavaError: An error occurred while calling o1442.load. : Failure to initialize configuration for storage account <storage-account>.dfs.core.windows.net: Invalid configuration value detected for fs.azure.account.keyInvalid configuration value detected for fs.azure.account.key at shaded.databricks.azurebfs.org.apache.hadoop.fs.azurebfs.services.SimpleKeyProvider.getStorageAccountKey(SimpleKeyProvider.java:52) at shaded.databricks.azurebfs.org.apache.hadoop.fs.azurebfs.AbfsConfiguration.getStorageAccountKey(AbfsConfiguration.java:666) at shaded.databricks.azurebfs.org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.initializeClient(AzureBlobFileSystemStore.java:2055) at shaded.databricks.azurebfs.org.apache.hadoop.fs.azurebfs.AzureBlobFileSystemStore.<init>(AzureBlobFileSystemStore.java:267) at shaded.databricks.azurebfs.org.apache.hadoop.fs.azurebfs.AzureBlobFileSystem.initialize(AzureBlobFileSystem.java:225) at com.databricks.common.filesystem.LokiFileSystem$.$anonfun$getLokiFS$1(LokiFileSystem.scala:63) at com.databricks.common.filesystem.Cache.getOrCompute(Cache.scala:38) at com.databricks.common.filesystem.LokiFileSystem$.getLokiFS(LokiFileSystem.scala:60) at com.databricks.common.filesystem.LokiFileSystem.initialize(LokiFileSystem.scala:86) at org.apache.hadoop.fs.FileSystem.createFileSystem(FileSystem.java:3469) at org.apache.hadoop.fs.FileSystem.get(FileSystem.java:537) at org.apache.hadoop.fs.Path.getFileSystem(Path.java:365) at com.databricks.sql.transaction.tahoe.DeltaValidation$.validateDeltaRead(DeltaValidation.scala:102) at org.apache.spark.sql.DataFrameReader.preprocessDeltaLoading(DataFrameReader.scala:280) at org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:329) at org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:240) at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method) at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62) at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43) at java.lang.reflect.Method.invoke(Method.java:498) at py4j.reflection.MethodInvoker.invoke(MethodInvoker.java:244) at py4j.reflection.ReflectionEngine.invoke(ReflectionEngine.java:380) at py4j.Gateway.invoke(Gateway.java:306) at py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:132) at py4j.commands.CallCommand.execute(CallCommand.java:79) at py4j.ClientServerConnection.waitForCommands(ClientServerConnection.java:195) at py4j.ClientServerConnection.run(ClientServerConnection.java:115) at java.lang.Thread.run(Thread.java:750) Caused by: Invalid configuration value detected for fs.azure.account.key at shaded.databricks.azurebfs.org.apache.hadoop.fs.azurebfs.diagnostics.ConfigurationBasicValidator.validate(ConfigurationBasicValidator.java:49) at shaded.databricks.azurebfs.org.apache.hadoop.fs.azurebfs.diagnostics.Base64StringConfigurationBasicValidator.validate(Base64StringConfigurationBasicValidator.java:40) at shaded.databricks.azurebfs.org.apache.hadoop.fs.azurebfs.services.SimpleKeyProvider.validateStorageAccountKey(SimpleKeyProvider.java:71) at shaded.databricks.azurebfs.org.apache.hadoop.fs.azurebfs.services.SimpleKeyProvider.getStorageAccountKey(SimpleKeyProvider.java:49)
核心差异与原因
- 执行上下文隔离:
%run直接在当前Notebook的Spark会话中执行代码,所有spark.conf配置会直接作用于当前会话;而dbutils.notebook.run()会启动独立的临时会话执行目标Notebook,执行完毕后临时会话销毁,配置不会传递回调用方的Spark会话。 - 配置生效范围:目标Notebook中设置的ADLS OAuth配置仅在自身临时会话内有效,调用方会话没有这些配置,导致读取Delta表时Spark尝试使用默认的密钥认证方式,因找不到有效密钥而报错。
解决方案
如果必须使用dbutils.notebook.run(),可通过以下方式传递配置:
- 返回配置参数:在
configure-storageNotebook末尾将配置信息以JSON格式返回,调用方接收后在自身会话中设置这些配置:
调用方代码:# 在configure-storage末尾添加 import json configs = { f"fs.azure.account.auth.type.{storage_account}.dfs.core.windows.net": "OAuth", f"fs.azure.account.oauth.provider.type.{storage_account}.dfs.core.windows.net": "org.apache.hadoop.fs.azurebfs.oauth2.ClientCredsTokenProvider", f"fs.azure.account.oauth2.client.id.{storage_account}.dfs.core.windows.net": client_id, f"fs.azure.account.oauth2.client.secret.{storage_account}.dfs.core.windows.net": client_secret, f"fs.azure.account.oauth2.client.endpoint.{storage_account}.dfs.core.windows.net": f"https://login.microsoftonline.com/{tenant_id}/oauth2/token" } dbutils.notebook.exit(json.dumps(configs))import json config_json = dbutils.notebook.run("../util/configure-storage", 300, { "storage_account": "storage-account-name", "tenant_id": "tenant-id-secret-name", "client_id": "client-id-secret-name", "client_secret": "client-secret-secret-name" }) configs = json.loads(config_json) for key, value in configs.items(): spark.conf.set(key, value) # 之后执行Delta表读取 file_location = "abfss://<storage-container>@<storage-account>.dfs.core.windows.net/<path-to-delta-table>" df = spark.read.format("delta").load(file_location) display(df) - 改用
%run:如果不需要异步执行或独立会话,%run是更直接的方式,配置会直接在当前会话生效。
内容的提问来源于stack exchange,提问作者Matthew Tatsch
相关产品推荐
相关产品推荐

