Apache NiFi任务无法自动终止,如何创建按需任务并判定任务成功?
Hey there! Let's break this down step by step since you're new to NiFi and dealing with that annoying non-stopping job issue. I've been there before, so I get how frustrating it is to see HDFS flooded with tiny files.
First off, the root cause here is likely that your MongoDB processor is set to run on a repeating schedule, rather than a one-time execution. Here's how to fix it:
- If you're using
QueryMongoDB(for full data pulls):- Open the processor's configuration, go to the Scheduling tab.
- Set the Scheduling Strategy to
Manual Trigger(this means it will only run when you explicitly trigger it, not on a loop). - Double-check your query (e.g.,
{}for full collection pull) and set a reasonableBatch Sizeto ensure all data is fetched in one or a few batches, rather than trickling indefinitely.
- Avoid circular flows: Make sure there's no connection from
PutHDFS(or any downstream processor) linking back to your MongoDB processor—this would create an infinite loop generating new files.
On-demand jobs in NiFi are all about triggering a one-time execution without persistent scheduling. Here are the most practical ways:
- Manual Trigger via UI:
- Select your entry processor (like
QueryMongoDBor even aGenerateFlowFileto kick off the workflow) - Right-click it and choose
Trigger Now—this runs the processor once, and it stops automatically once all data is processed. - For reusable on-demand workflows, save your MongoDB→HDFS flow as a Template. Each time you need to run the job, instantiate the template, trigger the entry processor, and delete the instance once done.
- Select your entry processor (like
- Trigger via REST API (for automation):
If you want to kick off the job from an external script or tool, use NiFi's REST API:
Replacecurl -X POST 'http://<nifi-host>:<port>/nifi-api/processors/<processor-id>/run-triggers' \ -H 'Content-Type: application/json' \ -d '{"count":1}'<nifi-host>,<port>, and<processor-id>with your actual values—this triggers a single run of the processor.
NiFi gives you multiple ways to verify your job completed successfully:
- Check Processor Metrics:
- Look at the
Successcount onQueryMongoDB—it should match the number of documents you expected to pull from MongoDB. - Confirm the
Successcount onPutHDFSmatches theSuccesscount from the MongoDB processor (or is proportional, if you're merging FlowFiles). - A
Failurecount of 0 on both processors means no errors occurred during the job.
- Look at the
- Validate FlowFile Attributes:
- Add a
LogAttributeprocessor afterPutHDFSto print key attributes to the NiFi log, like:mongodb.query.count(number of docs pulled from MongoDB)hdfs.file.path(path of the saved file in HDFS)file.size(size of the written file)
- You can also use
UpdateAttributeto tag successful FlowFiles with a custom attribute (e.g.,job_status=success) for easy tracking.
- Add a
- Post-Job Manual/Automated Checks:
- Manually verify HDFS: Run
hdfs dfs -ls <your-target-path>to confirm files exist, andhdfs dfs -cat <file-path> | headto check data format. - Add an
ExecuteStreamCommandprocessor afterPutHDFSto run a validation script—this script can check file counts/sizes in HDFS and send a success/failure notification (e.g., email, Slack).
- Manually verify HDFS: Run
To recap: Set your MongoDB processor to manual trigger for one-time runs, avoid circular flows, use processor metrics and FlowFile attributes to confirm success, and add validation steps if you need concrete proof data landed correctly.
内容的提问来源于stack exchange,提问作者Vivek Singh

