Partition Record Processor处理CSV数据报错:封装令牌与分隔符间含无效字符
Let's break down how to resolve this issue, since you can't modify the API's schema or output. First, let's diagnose the likely problem by comparing your two lines:
Looking at line 80 vs 81, the most probable culprits are either:
- A mismatched field count (even though it looks like both have the same number of fields, a hidden trailing comma or unescaped character might be throwing off the parser)
- The final empty field in line 81 (
"") conflicting with your schema's requirements (e.g., the schema expects a non-null value here) - A type mismatch (e.g., a field that the schema expects to be a string is coming through as an empty value that the parser can't resolve)
Here are actionable solutions you can implement in NiFi:
1. Tweak the CSV Reader Configuration (Quick Win)
First, check the CSV Reader settings used by your Partition Record Processor. Adjust these parameters to be more lenient with the API's output:
- Enable
Allow Trailing Delimiter: This lets the parser ignore extra commas at the end of lines, which fixes issues with empty trailing fields. - Set
Null Valueto"": Tells the parser to treat empty quoted strings as null values, which aligns with most schemas that allow nullable fields. - Enable
Trim Fields: Removes accidental whitespace from field values (this can fix unexpected type conversion errors). - Confirm
Quote Characteris set to": Ensures fields with spaces (like"abc travels and Cargo") are parsed correctly.
If this fixes the error, you're done! If not, move to the next step.
2. Preprocess the CSV to Fix Problematic Rows
If adjusting the reader isn't enough, add a preprocessing pipeline to clean the CSV before it reaches the Partition Record Processor:
Step 1: Split the CSV into Individual Rows
Use the SplitText processor to break the entire CSV file into single-line FlowFiles:
- Set
Line Split Countto1so each row becomes its own FlowFile. - Enable
Include Header Lineif your CSV has a header, so you can reattach it later.
Step 2: Fix the Problematic Line(s)
Use ReplaceText or UpdateRecord to correct line 81 (and any similar rows):
- If it's a missing field: Use a regex in
ReplaceTextto append the missing empty field to the end of the line. For example, if the line ends with0,instead of0,"", use the regex^(.*0,)$and replacement value$1"". - If it's a type mismatch: Use
UpdateRecordto set a default value for the problematic field. For example, if the final field expects a string, set its value to"N/A"when it's empty using the expression${field.value:isEmpty():ifElse("N/A", field.value)}.
Step 3: Merge the Rows Back into a CSV
Use MergeContent to combine the cleaned single-line FlowFiles back into a complete CSV:
- Set
Delimiter StrategytoTextandHeader Lineto your CSV's header (if you included it earlier). - Set
Footer Lineto empty, and ensure the delimiter is a newline character (\n).
3. Validate and Route Error Rows (For Robustness)
For a more resilient pipeline, add validation to catch and fix errors automatically:
- ValidateRecord: Configure it with the same schema and CSV Reader as your Partition Record Processor. Set
Invalid Record Routeto send bad rows to a separate relationship. - RouteOnAttribute: Use the
record_validation_errorsattribute to route rows based on their error type (e.g., field count mismatch vs type error). - Fix Errors: Use
UpdateRecordorReplaceTexton the error branch to correct the issue, then merge the fixed rows back with the valid ones before processing.
Specific Fix for Your Line 81
Looking at your line 81, the final field is an empty string (""), while line 80 has a value like "4:-1:10:0". If your schema expects a non-null value here, use UpdateRecord to set a default value for that field when it's empty. For example, if the field is named last_field, set its value to "DEFAULT" using the expression language.
内容的提问来源于stack exchange,提问作者prathik vijaykumar

