You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Azure Data Factory中高效获取SFTP全量文件同步至ADLS

解决Azure Data Factory Metadata活动获取SFTP全量文件(含根目录)的问题

问题分析

你用rootpath/**/**/**作为通配符路径无法捕获根目录文件,是因为多层**仅匹配子目录层级,根目录下的文件不在匹配范围内。以下是无需ForEach的高效实现方案:

步骤1:配置SFTP数据集

  • 创建/编辑SFTP数据集,将文件路径设为rootpath(直接指向根目录,不要加通配符)
  • 确保数据集的服务器地址、认证方式等连接配置正确

步骤2:配置Metadata活动

  1. 拖入Metadata活动,选择上述SFTP数据集作为数据源
  2. 在字段列表中勾选childItems(必选),若需基于时间筛选可同时勾选lastModified
  3. 开启递归选项(核心设置,会遍历rootpath下所有层级的子目录)
  4. 添加筛选条件:
    • 第一行:type → equals → File(仅获取文件,排除目录)
    • 第二行:lastModified → greater than or equal to → @pipeline().parameters.last_modified_timestamp(替换为你的时间参数)

步骤3:转换childItems格式(若需纯文件名)

默认childItems中的name会包含相对路径(如dir1/test2.csv),若要得到仅文件名的数组,可添加一个Set Variable活动:

  • 变量类型设为Array
  • 变量值使用以下表达式:
@map(activity('Get Metadata').output.childItems, item() => createObject('name', split(item().name, '/')[sub(length(split(item().name, '/')),1)], 'type', item().type))

该表达式会从带路径的文件名中提取纯文件名,生成你预期的数组结构。

验证输出

执行管道后,Set Variable的输出即为符合要求的childItems数组:

"childItems": [
        {
            "name": "test.csv",
            "type": "File"
        },
        {
            "name": "test1.csv",
            "type": "File"
        },
        {
            "name": "test2.csv",
            "type": "File"
        },
        {
            "name": "test3.zip",
            "type": "File"
        },
        {
            "name": "test4.csv",
            "type": "File"
        },
        {
            "name": "test5.csv",
            "type": "File"
        },
        {
            "name": "test6.zip",
            "type": "File"
        }
]

内容的提问来源于stack exchange,提问作者JasonM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 04:07:39