NiFi ExtractText处理器异常求助:数据无法流转至该处理器
Let's break down your NiFi issue step by step—sounds like you've got a couple of linked problems here: backlogged data piling up before SplitText, an unresponsive ExtractText processor, and the need to filter rows where the 18th field is BT, CV7, or CV30. Here's how to tackle each part:
First, let's get ExtractText up and running, since its unresponsiveness is likely contributing to the upstream backlog:
- Check validation errors: Hover over the ExtractText icon on the NiFi canvas. If there's a red exclamation mark, click it to see what's wrong. Common culprits include missing required properties (like leaving the
Regular ExpressionorDestinationfield blank) or invalid regex syntax (even a misplaced bracket can break it). - Verify scheduling and enablement: Make sure the processor is actually enabled (right-click > Enable). Double-check its scheduling strategy—if it's set to
Timer Drivenwith an overly long interval, orEvent Drivenbut not connected to receive flowfiles, it won't trigger. - Dig into NiFi logs: The
nifi-app.logfile will have specific error messages about why ExtractText isn't starting. Look for entries mentioning "ExtractText" to pinpoint the root cause (e.g., permission issues, invalid configuration values).
Once ExtractText is working, the backlog before SplitText should start clearing, but let's confirm the pipeline is set up correctly:
- Validate SplitText configuration: Ensure SplitText is properly splitting your input into individual rows. Check the
Split Strategy(set toSplit into linesif you're dealing with row-based data) andLine Split Count(usually 1 for one row per flowfile). If your input isn't being split into rows, downstream processors have nothing to work with. - Check connections: Verify the relationship from SplitText to ExtractText is set to the correct one (usually the
splitrelationship) and that ExtractText is configured to accept flowfiles from that connection. - Review backpressure settings: NiFi might be throttling upstream if downstream processors can't keep up. You can adjust backpressure thresholds in Controller Settings > Backpressure, but only do this after fixing the core issue with ExtractText—otherwise, you'll just push more data into a broken pipeline.
Once your pipeline is flowing, let's get the filtering right. ExtractText is great for extracting values, but RouteOnContent is more straightforward for conditional routing. Here are two approaches:
Option 1: Use RouteOnContent (Recommended)
This processor lets you filter rows directly without extra steps:
- Add a
RouteOnContentprocessor right after SplitText. - Set the
Matching StrategytoRegular Expression. - Create a new property (e.g.,
Keep Valid Rows) with this regex (adjust the delimiter,to match your actual field separator—use\tfor tabs):
The regex works by skipping the first 17 fields (each followed by a delimiter) then checking if the 18th field is one of your target values.^(?:[^,]+,){17}(BT|CV7|CV30), - Route the
matchedrelationship to your next processor, and sendunmatchedrows to a debug destination (likePutFileorLogAttribute) to verify filtering is working.
Option 2: Extract + RouteWithAttribute
If you need to capture the 18th field as an attribute first, use ExtractText followed by RouteOnAttribute:
- In ExtractText, set the regex to capture the 18th field:
^(?:[^,]+,){17}(BT|CV7|CV30) - Set
Destinationtoflowfile-attributeand name the attribute (e.g.,target_field). - Add a
RouteOnAttributeprocessor, create a property likeKeep Validwith the expression${target_field:isNotNull()}. - Route
matchedflowfiles to your desired destination, andunmatchedto debug.
内容的提问来源于stack exchange,提问作者omer

