You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HDFS指定文件夹叶子节点中小文件合并方案咨询

Hey there! Dealing with small files cluttering up HDFS leaf directories is such a common headache—especially when you’ve got a mix of text and Parquet files scattered across paths like /aaa/bbb/ccc/x/ and partitioned folders like /aaa/bbb/ccc/t/dt=xxxx-xx-xx/. Let me break down a few practical, efficient methods to merge all those sub-10MB files into single files per leaf directory.

1. HDFS Native Tools (Great for Text Files)

If you’re only dealing with text files, HDFS’s built-in commands are quick and straightforward. You can use a combination of hdfs dfs -cat and hdfs dfs -put to merge files directly in HDFS:

  • Merge a single leaf directory:

    # Merge all txt files in /aaa/bbb/ccc/x/ into one file
    hdfs dfs -cat /aaa/bbb/ccc/x/*.txt | hdfs dfs -put - /aaa/bbb/ccc/x/merged_all.txt
    
    # Verify the merged file size first, then delete the original small files
    hdfs dfs -rm /aaa/bbb/ccc/x/x*.txt
    
  • Batch process multiple leaf directories:
    Write a simple shell script to loop through all target directories. For example:

    #!/bin/bash
    HDFS_ROOT="/aaa/bbb/ccc"
    
    # Find all leaf directories containing txt files
    hdfs dfs -find $HDFS_ROOT -type f -name "*.txt" | sed 's/\/[^/]*$//' | sort | uniq | while read dir; do
        echo "Merging files in $dir..."
        hdfs dfs -cat "$dir"/*.txt | hdfs dfs -put - "$dir"/merged_all.txt
        # Optional: Delete original files after verification
        # hdfs dfs -rm "$dir"/*.txt
    done
    

2. Spark (Best for Mixed Text/Parquet & Large-Scale Jobs)

Spark is the go-to tool when you need to handle both text and Parquet files at scale, especially since Parquet can’t be merged like plain text (it’s a columnar storage format).

Merge Text Files with Spark (Python Example)

from pyspark.sql import SparkSession

spark = SparkSession.builder.appName("SmallTextMerger").getOrCreate()

# Read all text files recursively, preserving file paths
text_rdd = spark.sparkContext.wholeTextFiles(f"/aaa/bbb/ccc/**/*.txt")

# Group files by their parent directory and merge content
def merge_by_dir(iterator):
    current_dir = None
    merged_content = []
    for file_path, content in iterator:
        dir_path = "/".join(file_path.split("/")[:-1])
        if current_dir is None:
            current_dir = dir_path
        merged_content.append(content)
    if current_dir:
        # Write merged content to a single file in the same directory
        output_path = f"{current_dir}/merged_all.txt"
        # Overwrite existing file if present
        spark.sparkContext.parallelize(merged_content).saveAsTextFile(output_path)
        # Optional: Delete original small files here (add HDFS delete logic)

text_rdd.groupBy(lambda x: "/".join(x[0].split("/")[:-1])).foreachPartition(merge_by_dir)

spark.stop()

Merge Parquet Files with Spark (Python Example)

Parquet files require proper schema-aware merging. Use Spark SQL to read the files and rewrite them with controlled file sizes:

from pyspark.sql import SparkSession

spark = SparkSession.builder.appName("SmallParquetMerger").getOrCreate()

# Configure Spark to write files around 10MB (adjust bytes as needed)
spark.conf.set("spark.sql.files.maxPartitionBytes", 10 * 1024 * 1024)  # 10MB

# Read all Parquet files recursively (including partitioned folders)
df = spark.read.parquet(f"/aaa/bbb/ccc/**/*.parquet")

# Write back to HDFS, preserving partition structure if needed
# If your Parquet files use partitions like dt=xxxx-xx-xx, use partitionBy()
df.write.mode("overwrite") \
  .partitionBy("dt") \
  .parquet(f"/aaa/bbb/ccc/t/merged_parquet")

# After verification, delete the original part-* files in the dt partitions
# hdfs dfs -rm /aaa/bbb/ccc/t/dt=*/part-*

spark.stop()

3. Hive (Perfect for Parquet Partition Tables)

If your Parquet files are already part of a Hive partitioned table, Hive’s built-in merging settings make this a breeze:

# Set Hive parameters to control merge behavior
SET hive.merge.mapfiles = true;
SET hive.merge.mapredfiles = true;
SET hive.merge.size.per.task = 10485760;  # Target file size: 10MB
SET hive.merge.smallfiles.avgsize = 5242880;  # Trigger merge if avg file size <5MB

# Overwrite the target partition with merged files
INSERT OVERWRITE TABLE your_parquet_table PARTITION (dt='xxxx-xx-xx')
SELECT * FROM your_parquet_table WHERE dt='xxxx-xx-xx';

Critical Notes Before You Start

  • Backup first: Always copy the original directories to a backup location before deleting any files. Use hdfs dfs -cp /aaa/bbb/ccc /aaa/bbb/ccc_backup to be safe.
  • Permissions: Ensure you have read/write/delete permissions on all target HDFS paths.
  • Parquet caveat: Never merge Parquet files using text-based tools (like cat)—this will corrupt the file structure. Always use Spark, Hive, or other Parquet-aware tools.

内容的提问来源于stack exchange,提问作者Ssong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:40:36