如何在Azure Data Factory中高效获取SFTP全量文件同步至ADLS
解决Azure Data Factory Metadata活动获取SFTP全量文件(含根目录)的问题
问题分析
你用rootpath/**/**/**作为通配符路径无法捕获根目录文件,是因为多层**仅匹配子目录层级,根目录下的文件不在匹配范围内。以下是无需ForEach的高效实现方案:
步骤1:配置SFTP数据集
- 创建/编辑SFTP数据集,将文件路径设为
rootpath(直接指向根目录,不要加通配符) - 确保数据集的服务器地址、认证方式等连接配置正确
步骤2:配置Metadata活动
- 拖入Metadata活动,选择上述SFTP数据集作为数据源
- 在字段列表中勾选
childItems(必选),若需基于时间筛选可同时勾选lastModified - 开启递归选项(核心设置,会遍历
rootpath下所有层级的子目录) - 添加筛选条件:
- 第一行:
type→equals→File(仅获取文件,排除目录) - 第二行:
lastModified→greater than or equal to→@pipeline().parameters.last_modified_timestamp(替换为你的时间参数)
- 第一行:
步骤3:转换childItems格式(若需纯文件名)
默认childItems中的name会包含相对路径(如dir1/test2.csv),若要得到仅文件名的数组,可添加一个Set Variable活动:
- 变量类型设为
Array - 变量值使用以下表达式:
@map(activity('Get Metadata').output.childItems, item() => createObject('name', split(item().name, '/')[sub(length(split(item().name, '/')),1)], 'type', item().type))
该表达式会从带路径的文件名中提取纯文件名,生成你预期的数组结构。
验证输出
执行管道后,Set Variable的输出即为符合要求的childItems数组:
"childItems": [ { "name": "test.csv", "type": "File" }, { "name": "test1.csv", "type": "File" }, { "name": "test2.csv", "type": "File" }, { "name": "test3.zip", "type": "File" }, { "name": "test4.csv", "type": "File" }, { "name": "test5.csv", "type": "File" }, { "name": "test6.zip", "type": "File" } ]
内容的提问来源于stack exchange,提问作者JasonM
相关产品推荐
相关产品推荐

