如何在Databricks工作区用Python Notebook程序化读取其他Notebook内容
在Databricks同一工作区程序化读取其他Notebook内容的解决方案
问题场景
在Databricks同一工作区中,尝试通过Notebook程序化读取另一个指定Notebook内容时触发错误。已成功提取当前Notebook路径并列出目录下的Notebook:
import os notebook_path = os.path.dirname(dbutils.entry_point.getDbutils().notebook().getContext().notebookPath().getOrElse(None)) for f in dbutils.fs.ls(f"file:/Workspace{notebook_path}"): print(f)
示例输出:
FileInfo(path='file:/Workspace/Users/***/Notebook1', name='Notebook1', size=0, modificationTime=0)
但使用Python原生open函数读取时出现以下错误:
with open(f"/Workspace{notebook_path}/Notebook1", "r") as _f: c = _f.read()
错误信息:
OSError: [Errno 95] Operation not supported File /databricks/python/lib/python3.10/site-packages/IPython/core/interactiveshell.py:282, in _modified_open(file, *args, **kwargs) 275 if file in {0, 1, 2}: 276 raise ValueError( 277 f"IPython won't let you open fd={file} by default " 278 "as it is likely to crash IPython. If you know what you are doing, " 279 "you can use builtins' open." 280 ) --> 282 return io_open(file, *args, **kwargs)
原因
Databricks的/Workspace路径属于虚拟文件系统(DBFS的Workspace挂载点),Python原生的文件操作函数(如open)无法直接访问该路径,因此会触发"Operation not supported"错误。
解决方案
方法1:复制到本地临时目录后读取
将目标Notebook从Workspace复制到集群的本地临时目录(如/tmp),再用open读取内容:
import os # 目标Notebook的Workspace完整路径 target_notebook = f"/Workspace{notebook_path}/Notebook1" # 本地临时存储路径 local_path = "/tmp/target_notebook" # 从Workspace复制到本地 dbutils.fs.cp(target_notebook, f"file:{local_path}", recurse=True) # 读取内容 with open(local_path, "r") as f: notebook_content = f.read() print(notebook_content) # 清理临时文件(可选) os.remove(local_path)
方法2:调用Databricks Workspace API
适合批量读取或需要更灵活控制的场景,通过API直接获取Notebook源代码:
import requests import base64 # 获取工作区API地址和认证令牌 workspace_url = dbutils.entry_point.getDbutils().notebook().getContext().apiUrl().getOrElse(None) auth_token = dbutils.entry_point.getDbutils().notebook().getContext().apiToken().getOrElse(None) # 目标Notebook的完整路径(格式:/Users/xxx/Notebook1) notebook_path_full = f"{notebook_path}/Notebook1" # 构建API请求 api_endpoint = f"{workspace_url}/api/2.0/workspace/export" headers = {"Authorization": f"Bearer {auth_token}"} params = { "path": notebook_path_full, "format": "SOURCE" # 输出格式:SOURCE(源代码)、HTML、JUPYTER、DBC } # 发送请求并解析结果 response = requests.get(api_endpoint, headers=headers, params=params) if response.status_code == 200: # API返回的内容是base64编码,需要解码 encoded_content = response.json()["content"] decoded_content = base64.b64decode(encoded_content).decode("utf-8") print(decoded_content) else: print(f"读取失败:{response.text}")
内容的提问来源于stack exchange,提问作者evg
相关产品推荐
相关产品推荐

