You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NiFi ExtractText处理器异常求助:数据无法流转至该处理器

Let's break down your NiFi issue step by step—sounds like you've got a couple of linked problems here: backlogged data piling up before SplitText, an unresponsive ExtractText processor, and the need to filter rows where the 18th field is BT, CV7, or CV30. Here's how to tackle each part:

Fix the ExtractText Processor Stagnation

First, let's get ExtractText up and running, since its unresponsiveness is likely contributing to the upstream backlog:

  • Check validation errors: Hover over the ExtractText icon on the NiFi canvas. If there's a red exclamation mark, click it to see what's wrong. Common culprits include missing required properties (like leaving the Regular Expression or Destination field blank) or invalid regex syntax (even a misplaced bracket can break it).
  • Verify scheduling and enablement: Make sure the processor is actually enabled (right-click > Enable). Double-check its scheduling strategy—if it's set to Timer Driven with an overly long interval, or Event Driven but not connected to receive flowfiles, it won't trigger.
  • Dig into NiFi logs: The nifi-app.log file will have specific error messages about why ExtractText isn't starting. Look for entries mentioning "ExtractText" to pinpoint the root cause (e.g., permission issues, invalid configuration values).
Resolve SplitText Upstream Backlog

Once ExtractText is working, the backlog before SplitText should start clearing, but let's confirm the pipeline is set up correctly:

  • Validate SplitText configuration: Ensure SplitText is properly splitting your input into individual rows. Check the Split Strategy (set to Split into lines if you're dealing with row-based data) and Line Split Count (usually 1 for one row per flowfile). If your input isn't being split into rows, downstream processors have nothing to work with.
  • Check connections: Verify the relationship from SplitText to ExtractText is set to the correct one (usually the split relationship) and that ExtractText is configured to accept flowfiles from that connection.
  • Review backpressure settings: NiFi might be throttling upstream if downstream processors can't keep up. You can adjust backpressure thresholds in Controller Settings > Backpressure, but only do this after fixing the core issue with ExtractText—otherwise, you'll just push more data into a broken pipeline.
Filter Rows by the 18th Field

Once your pipeline is flowing, let's get the filtering right. ExtractText is great for extracting values, but RouteOnContent is more straightforward for conditional routing. Here are two approaches:

This processor lets you filter rows directly without extra steps:

  1. Add a RouteOnContent processor right after SplitText.
  2. Set the Matching Strategy to Regular Expression.
  3. Create a new property (e.g., Keep Valid Rows) with this regex (adjust the delimiter , to match your actual field separator—use \t for tabs):
    ^(?:[^,]+,){17}(BT|CV7|CV30),
    
    The regex works by skipping the first 17 fields (each followed by a delimiter) then checking if the 18th field is one of your target values.
  4. Route the matched relationship to your next processor, and send unmatched rows to a debug destination (like PutFile or LogAttribute) to verify filtering is working.

Option 2: Extract + RouteWithAttribute

If you need to capture the 18th field as an attribute first, use ExtractText followed by RouteOnAttribute:

  1. In ExtractText, set the regex to capture the 18th field:
    ^(?:[^,]+,){17}(BT|CV7|CV30)
    
  2. Set Destination to flowfile-attribute and name the attribute (e.g., target_field).
  3. Add a RouteOnAttribute processor, create a property like Keep Valid with the expression ${target_field:isNotNull()}.
  4. Route matched flowfiles to your desired destination, and unmatched to debug.

内容的提问来源于stack exchange,提问作者omer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:16:46