如何从NiFi的FetchS3Object文件名属性提取倒数第二级路径文本?
Great question! Here are a couple of straightforward ways to extract that second-last path segment in NiFi, depending on your preference for no-code or scripting approaches:
Method 1: Use UpdateAttribute with NiFi Expression Language (No Scripting)
This is the simplest approach for most cases:
- Add an
UpdateAttributeprocessor to your flow. - Create a new user-defined attribute (e.g.,
second_last_level). - Set the attribute's value to this EL snippet:
${filename:split('/'):get(${filename:split('/'):size() - 3})} - How it works:
filename:split('/')breaks your path into an array of segments (like["root1", "subobject", ..., "text.csv"]).size()gives the total number of segments. Subtracting 3 skips the filename itself and its immediate parent directory, landing you on the second-last level.get(...)pulls that target segment from the array.
Method 2: Regex Replacement in UpdateAttribute
If you prefer regex over array manipulation, use replaceAll:
- In the same
UpdateAttributeprocessor, setsecond_last_levelto:${filename:replaceAll('^.*\\/([^\\/]+)\\/[^\\/]+\\/[^\\/]+$', '$1')} - How it works:
- The regex matches the entire path, capturing the segment that sits between the penultimate directory and the rest of the path.
$1references the captured group, which is exactly your desiredpath2segment.
Method 3: ExecuteScript for Flexible Logic
If you need error handling for shorter paths or want more control, use ExecuteScript with Groovy:
- Add an
ExecuteScriptprocessor, set the Script Language to Groovy. - Paste this code:
def filename = flowFile.getAttribute('filename') if (filename) { def parts = filename.split('/') if (parts.length >= 3) { // Ensure there are enough segments to extract from def secondLastLevel = parts[parts.length - 3] flowFile.setAttribute('second_last_level', secondLastLevel) } else { // Handle edge cases (e.g., short paths) by setting a default value flowFile.setAttribute('second_last_level', 'N/A') } } return flowFile - This adds safeguards for cases where your S3 keys might have fewer directory levels than expected.
All these methods work directly with the filename attribute you described (since it already excludes the bucket name from the full key path).
内容的提问来源于stack exchange,提问作者data_addict
相关产品推荐
相关产品推荐

