You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Apache NiFi任务无法自动终止,如何创建按需任务并判定任务成功?

Hey there! Let's break this down step by step since you're new to NiFi and dealing with that annoying non-stopping job issue. I've been there before, so I get how frustrating it is to see HDFS flooded with tiny files.

1. Fixing the Non-Stop Job Issue (Stopping After Full Import)

First off, the root cause here is likely that your MongoDB processor is set to run on a repeating schedule, rather than a one-time execution. Here's how to fix it:

  • If you're using QueryMongoDB (for full data pulls):
    1. Open the processor's configuration, go to the Scheduling tab.
    2. Set the Scheduling Strategy to Manual Trigger (this means it will only run when you explicitly trigger it, not on a loop).
    3. Double-check your query (e.g., {} for full collection pull) and set a reasonable Batch Size to ensure all data is fetched in one or a few batches, rather than trickling indefinitely.
  • Avoid circular flows: Make sure there's no connection from PutHDFS (or any downstream processor) linking back to your MongoDB processor—this would create an infinite loop generating new files.
2. Creating an On-Demand Job in NiFi

On-demand jobs in NiFi are all about triggering a one-time execution without persistent scheduling. Here are the most practical ways:

  • Manual Trigger via UI:
    • Select your entry processor (like QueryMongoDB or even a GenerateFlowFile to kick off the workflow)
    • Right-click it and choose Trigger Now—this runs the processor once, and it stops automatically once all data is processed.
    • For reusable on-demand workflows, save your MongoDB→HDFS flow as a Template. Each time you need to run the job, instantiate the template, trigger the entry processor, and delete the instance once done.
  • Trigger via REST API (for automation):
    If you want to kick off the job from an external script or tool, use NiFi's REST API:
    curl -X POST 'http://<nifi-host>:<port>/nifi-api/processors/<processor-id>/run-triggers' \
         -H 'Content-Type: application/json' \
         -d '{"count":1}'
    
    Replace <nifi-host>, <port>, and <processor-id> with your actual values—this triggers a single run of the processor.
3. Determining if the Job Succeeded

NiFi gives you multiple ways to verify your job completed successfully:

  • Check Processor Metrics:
    • Look at the Success count on QueryMongoDB—it should match the number of documents you expected to pull from MongoDB.
    • Confirm the Success count on PutHDFS matches the Success count from the MongoDB processor (or is proportional, if you're merging FlowFiles).
    • A Failure count of 0 on both processors means no errors occurred during the job.
  • Validate FlowFile Attributes:
    • Add a LogAttribute processor after PutHDFS to print key attributes to the NiFi log, like:
      • mongodb.query.count (number of docs pulled from MongoDB)
      • hdfs.file.path (path of the saved file in HDFS)
      • file.size (size of the written file)
    • You can also use UpdateAttribute to tag successful FlowFiles with a custom attribute (e.g., job_status=success) for easy tracking.
  • Post-Job Manual/Automated Checks:
    • Manually verify HDFS: Run hdfs dfs -ls <your-target-path> to confirm files exist, and hdfs dfs -cat <file-path> | head to check data format.
    • Add an ExecuteStreamCommand processor after PutHDFS to run a validation script—this script can check file counts/sizes in HDFS and send a success/failure notification (e.g., email, Slack).

To recap: Set your MongoDB processor to manual trigger for one-time runs, avoid circular flows, use processor metrics and FlowFile attributes to confirm success, and add validation steps if you need concrete proof data landed correctly.

内容的提问来源于stack exchange,提问作者Vivek Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:14:19