You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flume技术疑问:HDFS中单个文件存在两条消息的问题

Fix for Multiple IBM MQ Messages in Single HDFS File via Apache Flume

Hey there! Let’s work through getting each IBM MQ message into its own HDFS file using Apache Flume. The issue you’re facing almost always ties back to how Flume’s HDFS sink handles batching and file rolling. Here are the actionable fixes you can implement:

1. Configure the HDFS Sink for Single-Message Files

Your current config doesn’t include the HDFS sink section, so let’s add and adjust those settings to ensure one message per file:

# Sink definition
u.sinks.k1.type=hdfs
u.sinks.k1.hdfs.path=hdfs://your-nn-address:port/your/target/directory
u.sinks.k1.hdfs.filePrefix=mq_msg
u.sinks.k1.hdfs.fileSuffix=.log

# Disable time-based file rolling (we'll use message count instead)
u.sinks.k1.hdfs.rollInterval=0

# Roll to a new file after every 1 message
u.sinks.k1.hdfs.rollCount=1

# Write one message at a time to HDFS (no batching)
u.sinks.k1.hdfs.batchSize=1

# Ensure files are flushed immediately (no in-memory caching delays)
u.sinks.k1.hdfs.callTimeout=10000
u.sinks.k1.hdfs.writeFormat=Text

Why these settings work:

  • rollCount=1: Directs Flume to create a new HDFS file as soon as one message is written.
  • rollInterval=0: Turns off automatic time-based rolling, so messages aren’t grouped into the same file if they arrive within the default 30-second window.
  • batchSize=1: Ensures the sink writes each message to HDFS right away instead of waiting to accumulate a batch.

2. Tweak the JMS Source to Fetch One Message at a Time

If your Flume JMS source is pulling multiple messages in one batch, that can also lead to merged files. Add the batchSize parameter to your source config and set it to 1:

u.sources.s1.type=jms
u.sources.s1.initialContextFactory=ABC
u.sources.s1.connectionFactory=<my connection factory>
u.sources.s1.providerURL=ABC
u.sources.s1.destinationName=r1
u.sources.s1.destinationType=QUEUE
# Fetch only one message per batch from IBM MQ
u.sources.s1.batchSize=1

This ensures the source pulls just a single message from the queue at a time, so your file channel and sink never have multiple messages pending to write to the same file.

3. Quick Channel Configuration Check

Your file channel’s transactionCapacity=10000 is fine for throughput, but since we’re targeting single-message files, just make sure the source’s batch size doesn’t exceed this value (which it won’t with batchSize=1). No changes needed here unless you adjust batch sizes later.

Final Steps

  • Restart Flume after updating the config to apply all changes.
  • Check Flume’s logs for any errors related to the sink or source settings.
  • Inspect the HDFS files to confirm each now contains only one message.

内容的提问来源于stack exchange,提问作者Rishu S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:24:59