Hadoop 3纠删码功能配置方法及默认是否仅用于存储冷文件?
Great question! Let's break down both your configuration query and the usage scenario question in detail:
一、如何配置Hadoop 3的纠删码功能
Hadoop 3 supports multiple erasure coding schemes (like Reed-Solomon (10,4), RS(6,3), etc.), and the configuration mainly involves these steps:
Enable erasure coding support
Hadoop 3.x has built-in support for erasure coding by default, but you can confirm and enforce it by modifying thehdfs-site.xmlconfiguration file:<property> <name>dfs.ec.enabled</name> <value>true</value> </property>Define erasure coding policies
You can use built-in policies or customize your own. First, check which policies are supported by your cluster:hdfs ec -listPoliciesIf your target policy isn't enabled, turn it on with this command:
hdfs ec -enablePolicy -policy RS-10-4-1024kBind the policy to a directory
Erasure coding is configured at the directory level. Assign your chosen policy to the target directory:hdfs ec -setPolicy -path /path/to/ec-target-dir -policy RS-10-4-1024kAll new files created in this directory will automatically use the erasure coding policy.
Verify the configuration
Check the policy bound to a directory:hdfs ec -getPolicy -path /path/to/ec-target-dirYou can also inspect the coding details of a specific file:
hdfs fsck /path/to/ec-file -files -blocks -locations
二、Is erasure coding only used for cold files by default?
By default, Hadoop 3 does not force erasure coding to be used only for cold files — erasure coding policies take effect when you manually bind them to directories, so you can apply them to any directory, whether it stores cold or hot data.
That said, erasure coding is traditionally recommended for cold files due to performance considerations:
- Data reconstruction with erasure coding requires calculating additional parity blocks, which brings extra CPU and network overhead (reconstruction often involves reading multiple remote blocks)
- Hot data requires frequent read/write operations, and the default 3-replica replication strategy has lower latency, making it more suitable for high-access-frequency data
So in practice, erasure coding is commonly used for cold file storage to save storage costs without impacting hot data performance. But this is a best practice, not a mandatory restriction from Hadoop.
Bonus: Storage advantages of erasure coding
As you mentioned, erasure coding drastically reduces storage overhead compared to replication:
- The standard Reed-Solomon (10,4) scheme only incurs 1.4x storage overhead (10 data blocks + 4 parity blocks)
- In contrast, HDFS's default 3-replica replication strategy has 3x overhead, creating a significant difference in storage costs.
This is the core reason Hadoop 3 introduced erasure coding — to address the high storage costs of large-scale clusters.
内容的提问来源于stack exchange,提问作者Nitesh Gupta

