You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从NiFi的FetchS3Object文件名属性提取倒数第二级路径文本?

Great question! Here are a couple of straightforward ways to extract that second-last path segment in NiFi, depending on your preference for no-code or scripting approaches:

Method 1: Use UpdateAttribute with NiFi Expression Language (No Scripting)

This is the simplest approach for most cases:

  • Add an UpdateAttribute processor to your flow.
  • Create a new user-defined attribute (e.g., second_last_level).
  • Set the attribute's value to this EL snippet:
    ${filename:split('/'):get(${filename:split('/'):size() - 3})}
    
  • How it works:
    • filename:split('/') breaks your path into an array of segments (like ["root1", "subobject", ..., "text.csv"]).
    • size() gives the total number of segments. Subtracting 3 skips the filename itself and its immediate parent directory, landing you on the second-last level.
    • get(...) pulls that target segment from the array.

Method 2: Regex Replacement in UpdateAttribute

If you prefer regex over array manipulation, use replaceAll:

  • In the same UpdateAttribute processor, set second_last_level to:
    ${filename:replaceAll('^.*\\/([^\\/]+)\\/[^\\/]+\\/[^\\/]+$', '$1')}
    
  • How it works:
    • The regex matches the entire path, capturing the segment that sits between the penultimate directory and the rest of the path.
    • $1 references the captured group, which is exactly your desired path2 segment.

Method 3: ExecuteScript for Flexible Logic

If you need error handling for shorter paths or want more control, use ExecuteScript with Groovy:

  • Add an ExecuteScript processor, set the Script Language to Groovy.
  • Paste this code:
    def filename = flowFile.getAttribute('filename')
    if (filename) {
        def parts = filename.split('/')
        if (parts.length >= 3) { // Ensure there are enough segments to extract from
            def secondLastLevel = parts[parts.length - 3]
            flowFile.setAttribute('second_last_level', secondLastLevel)
        } else {
            // Handle edge cases (e.g., short paths) by setting a default value
            flowFile.setAttribute('second_last_level', 'N/A')
        }
    }
    return flowFile
    
  • This adds safeguards for cases where your S3 keys might have fewer directory levels than expected.

All these methods work directly with the filename attribute you described (since it already excludes the bucket name from the full key path).

内容的提问来源于stack exchange,提问作者data_addict

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:22:51