Hadoop fs -du命令双输出值含义及路径真实大小咨询
Understanding Hadoop
fs -du Output Values Let me break down what those two numbers mean clearly—since I get why the official docs can feel a bit vague on this:
- The first value (
5.9 Gin your output) is the actual, un-replicated size of your data in that path. This is the "true" size you’d see if you copied those files out of HDFS to a local filesystem; it’s just the raw size of the files themselves, no replication overhead included. - The second value (
17.8 G) is the total disk space consumed across the entire HDFS cluster for that data, including Hadoop’s replication factor. By default, HDFS uses a replication factor of 3, which explains why this number is roughly 3x the first one (5.9 * 3 ≈ 17.7, which matches your output almost exactly).
To answer your core question straight up: The real size of your data at /project/crm/warehouse/n98770_patron_1 is 5.9 GB. The 17.8 GB is just the total space taken up across all datanodes because of replication.
If you want to double-check, you can look up your cluster’s default replication factor (usually set in hdfs-site.xml under the dfs.replication property) and you’ll see it lines up perfectly with the ratio between the two numbers.
内容的提问来源于stack exchange,提问作者Carmen Pérez Carrillo
相关产品推荐
相关产品推荐

