Azure Databricks:从DBFS批量迁移文件/文件夹至用户工作区的方案
从DBFS迁移文件/文件夹到Azure Databricks用户工作区的可行方法
下面是几种高效的迁移方案,无需逐个手动上传:
1. 使用Databricks CLI批量复制
这是最直接的方法,适合本地已配置好CLI的场景,支持递归复制文件夹,大文件可直接后端传输(无需本地中转)。
- 先确认已安装Databricks CLI并完成认证配置(执行
databricks configure命令) - 执行复制命令:
# 复制单个文件 databricks fs cp dbfs:/path/to/source-file.py /Workspace/Users/your-email@domain.com/target-file.py # 递归复制整个文件夹 databricks fs cp dbfs:/path/to/source-folder /Workspace/Users/your-email@domain.com/target-folder --recursive
2. 在Databricks Notebook中用dbutils命令复制
直接在Databricks环境内操作,无需本地工具,适合临时迁移需求。
打开任意Python/Scala Notebook,运行以下代码:
# Python示例:递归复制文件夹到工作区 source_path = "dbfs:/path/to/source-folder" target_path = "/Workspace/Users/your-email@domain.com/target-folder" # 可选:先检查目标路径是否存在,避免覆盖 try: dbutils.fs.ls(target_path) print("目标路径已存在,请确认后再操作") except Exception: dbutils.fs.cp(source_path, target_path, recurse=True) print("复制完成")// Scala示例:递归复制文件夹到工作区 val sourcePath = "dbfs:/path/to/source-folder" val targetPath = "/Workspace/Users/your-email@domain.com/target-folder" if (dbutils.fs.ls(targetPath).isEmpty) { dbutils.fs.cp(sourcePath, targetPath, recurse = true) println("复制完成") } else { println("目标路径已存在,请确认后再操作") }
3. 使用Databricks REST API实现自动化迁移
适合批量、自动化迁移大量文件的场景,通过API遍历DBFS文件并导入到工作区。
- 示例Python代码(需提前获取Databricks个人访问令牌):
import requests # 配置基础信息 databricks_token = "your-personal-access-token" workspace_url = "https://<your-workspace-name>.databricks.com" headers = {"Authorization": f"Bearer {databricks_token}"} # 遍历DBFS目录下的所有文件和子文件夹 def list_dbfs_items(dbfs_path): resp = requests.get( f"{workspace_url}/api/2.0/dbfs/list", headers=headers, params={"path": dbfs_path} ) return resp.json().get("files", []) # 创建工作区文件夹 def create_workspace_folder(folder_path): requests.post( f"{workspace_url}/api/2.0/workspace/mkdirs", headers=headers, json={"path": folder_path} ) # 将DBFS文件导入到工作区 def import_file_to_workspace(dbfs_file_path, workspace_file_path): # 读取DBFS文件内容 file_resp = requests.get( f"{workspace_url}/api/2.0/dbfs/read", headers=headers, params={"path": dbfs_file_path} ) file_content = file_resp.content.decode("utf-8") # 导入到工作区 requests.post( f"{workspace_url}/api/2.0/workspace/import", headers=headers, json={ "path": workspace_file_path, "format": "AUTO", "content": file_content, "overwrite": True } ) # 递归复制DBFS文件夹到工作区 def copy_dbfs_folder_to_workspace(source_dbfs, target_workspace): create_workspace_folder(target_workspace) for item in list_dbfs_items(source_dbfs): if item["is_dir"]: sub_target = f"{target_workspace}/{item['path'].split('/')[-1]}" copy_dbfs_folder_to_workspace(item["path"], sub_target) else: file_name = item["path"].split('/')[-1] workspace_file_path = f"{target_workspace}/{file_name}" import_file_to_workspace(item["path"], workspace_file_path) # 执行迁移 copy_dbfs_folder_to_workspace( "dbfs:/path/to/source-folder", "/Workspace/Users/your-email@domain.com/target-folder" )
内容的提问来源于stack exchange,提问作者Programmer
相关产品推荐
相关产品推荐

