如何用Python或Pandas获取HDFS文件夹下的所有文件列表?
用Python结合subprocess获取HDFS文件列表并存入Pandas DataFrame
核心思路
通过subprocess调用HDFS原生命令hdfs dfs -ls,利用命令参数简化输出或手动解析文本过滤冗余信息,提取纯文件路径后转存为Pandas DataFrame。
实现方案
方案1:用-C参数直接获取纯文件路径(推荐)
hdfs dfs -ls默认会输出权限、所有者、文件大小等冗余信息,加上-C参数后可直接返回仅包含文件路径的结果,省去复杂解析步骤。如果需要递归遍历子文件夹,追加-R参数即可。
import subprocess import pandas as pd def get_hdfs_files(hdfs_path): try: # 执行HDFS命令,-C只返回文件路径,-R可选(递归遍历) output = subprocess.check_output( ["hdfs", "dfs", "-ls", "-C", hdfs_path], stderr=subprocess.STDOUT, text=True ) except subprocess.CalledProcessError as e: print(f"HDFS命令执行失败: {e.output}") return pd.DataFrame() # 分割输出行,过滤空行 file_paths = [line.strip() for line in output.split('\n') if line.strip()] # 转成DataFrame return pd.DataFrame(file_paths, columns=["hdfs_file_path"]) # 使用示例 target_dir = "/user/your_target_dir/" file_df = get_hdfs_files(target_dir) print(file_df)
方案2:手动解析默认ls输出(适配无-C参数的环境)
如果你的HDFS环境不支持-C参数,可通过分割默认输出的文本字段,提取最后一列的文件路径:
import subprocess import pandas as pd def parse_hdfs_files(hdfs_path): try: output = subprocess.check_output( ["hdfs", "dfs", "-ls", hdfs_path], stderr=subprocess.STDOUT, text=True ) except subprocess.CalledProcessError as e: print(f"HDFS命令执行失败: {e.output}") return pd.DataFrame() file_paths = [] for line in output.split('\n'): line = line.strip() # 跳过开头的"Found X items"提示行和空行 if not line or line.startswith("Found"): continue # 默认输出的最后一个字段是文件路径 parts = line.split() if len(parts) >= 8: file_paths.append(parts[-1]) return pd.DataFrame(file_paths, columns=["hdfs_file_path"])
关键说明
- 错误处理:捕获
CalledProcessError可处理路径不存在、权限不足等异常,避免程序直接崩溃。 - 文本模式:
text=True让命令输出直接返回字符串,无需手动解码字节流。 - 递归遍历:只需在命令参数中加入
-R,即可获取目标文件夹下所有子目录的文件。
内容的提问来源于stack exchange,提问作者Tom Bellmer
相关产品推荐
相关产品推荐

