HDFS指定文件夹叶子节点中小文件合并方案咨询
Hey there! Dealing with small files cluttering up HDFS leaf directories is such a common headache—especially when you’ve got a mix of text and Parquet files scattered across paths like /aaa/bbb/ccc/x/ and partitioned folders like /aaa/bbb/ccc/t/dt=xxxx-xx-xx/. Let me break down a few practical, efficient methods to merge all those sub-10MB files into single files per leaf directory.
1. HDFS Native Tools (Great for Text Files)
If you’re only dealing with text files, HDFS’s built-in commands are quick and straightforward. You can use a combination of hdfs dfs -cat and hdfs dfs -put to merge files directly in HDFS:
Merge a single leaf directory:
# Merge all txt files in /aaa/bbb/ccc/x/ into one file hdfs dfs -cat /aaa/bbb/ccc/x/*.txt | hdfs dfs -put - /aaa/bbb/ccc/x/merged_all.txt # Verify the merged file size first, then delete the original small files hdfs dfs -rm /aaa/bbb/ccc/x/x*.txtBatch process multiple leaf directories:
Write a simple shell script to loop through all target directories. For example:#!/bin/bash HDFS_ROOT="/aaa/bbb/ccc" # Find all leaf directories containing txt files hdfs dfs -find $HDFS_ROOT -type f -name "*.txt" | sed 's/\/[^/]*$//' | sort | uniq | while read dir; do echo "Merging files in $dir..." hdfs dfs -cat "$dir"/*.txt | hdfs dfs -put - "$dir"/merged_all.txt # Optional: Delete original files after verification # hdfs dfs -rm "$dir"/*.txt done
2. Spark (Best for Mixed Text/Parquet & Large-Scale Jobs)
Spark is the go-to tool when you need to handle both text and Parquet files at scale, especially since Parquet can’t be merged like plain text (it’s a columnar storage format).
Merge Text Files with Spark (Python Example)
from pyspark.sql import SparkSession spark = SparkSession.builder.appName("SmallTextMerger").getOrCreate() # Read all text files recursively, preserving file paths text_rdd = spark.sparkContext.wholeTextFiles(f"/aaa/bbb/ccc/**/*.txt") # Group files by their parent directory and merge content def merge_by_dir(iterator): current_dir = None merged_content = [] for file_path, content in iterator: dir_path = "/".join(file_path.split("/")[:-1]) if current_dir is None: current_dir = dir_path merged_content.append(content) if current_dir: # Write merged content to a single file in the same directory output_path = f"{current_dir}/merged_all.txt" # Overwrite existing file if present spark.sparkContext.parallelize(merged_content).saveAsTextFile(output_path) # Optional: Delete original small files here (add HDFS delete logic) text_rdd.groupBy(lambda x: "/".join(x[0].split("/")[:-1])).foreachPartition(merge_by_dir) spark.stop()
Merge Parquet Files with Spark (Python Example)
Parquet files require proper schema-aware merging. Use Spark SQL to read the files and rewrite them with controlled file sizes:
from pyspark.sql import SparkSession spark = SparkSession.builder.appName("SmallParquetMerger").getOrCreate() # Configure Spark to write files around 10MB (adjust bytes as needed) spark.conf.set("spark.sql.files.maxPartitionBytes", 10 * 1024 * 1024) # 10MB # Read all Parquet files recursively (including partitioned folders) df = spark.read.parquet(f"/aaa/bbb/ccc/**/*.parquet") # Write back to HDFS, preserving partition structure if needed # If your Parquet files use partitions like dt=xxxx-xx-xx, use partitionBy() df.write.mode("overwrite") \ .partitionBy("dt") \ .parquet(f"/aaa/bbb/ccc/t/merged_parquet") # After verification, delete the original part-* files in the dt partitions # hdfs dfs -rm /aaa/bbb/ccc/t/dt=*/part-* spark.stop()
3. Hive (Perfect for Parquet Partition Tables)
If your Parquet files are already part of a Hive partitioned table, Hive’s built-in merging settings make this a breeze:
# Set Hive parameters to control merge behavior SET hive.merge.mapfiles = true; SET hive.merge.mapredfiles = true; SET hive.merge.size.per.task = 10485760; # Target file size: 10MB SET hive.merge.smallfiles.avgsize = 5242880; # Trigger merge if avg file size <5MB # Overwrite the target partition with merged files INSERT OVERWRITE TABLE your_parquet_table PARTITION (dt='xxxx-xx-xx') SELECT * FROM your_parquet_table WHERE dt='xxxx-xx-xx';
Critical Notes Before You Start
- Backup first: Always copy the original directories to a backup location before deleting any files. Use
hdfs dfs -cp /aaa/bbb/ccc /aaa/bbb/ccc_backupto be safe. - Permissions: Ensure you have read/write/delete permissions on all target HDFS paths.
- Parquet caveat: Never merge Parquet files using text-based tools (like
cat)—this will corrupt the file structure. Always use Spark, Hive, or other Parquet-aware tools.
内容的提问来源于stack exchange,提问作者Ssong

