You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手遇HDFS文件读取IOError:文件或目录不存在

解决Python读取HDFS文件时的IOError问题

Hey there! I see you're a Python newbie running into a tricky issue with reading HDFS files—let's break this down and fix it.

为什么会报错?

Python的内置open()函数只能处理本地文件系统的文件,它根本不认识hdfs:///这种HDFS协议的路径。所以哪怕你用hdfs dfs -ls能看到文件,open()还是会把它当成本地路径去找,自然就找不到啦,这就是你看到IOError: [Errno 2] No such file or directory的原因。

三种解决方案,按需选择:

1. 用专门的HDFS Python库(最推荐!)

如果需要经常和HDFS交互,安装hdfs库是最靠谱的方式:
首先在终端安装库:

pip install hdfs

然后修改你的代码,用这个库连接HDFS并读取文件:

from hdfs import InsecureClient
import json

# 替换成你的HDFS namenode地址和用户名
client = InsecureClient('http://your-namenode-ip:50070', user='your-username')

# 读取HDFS上的文件
with client.read('/data/testdata.json') as data_file:
    file_content = data_file.read()
    json_data = json.loads(file_content)
    # 这里就可以处理你的数据啦

注意:如果你的HDFS集群开启了Kerberos认证,需要用KerberosClient代替InsecureClient,可以根据库的文档调整配置。

2. 先把HDFS文件下载到本地再读取

如果只是临时处理小文件,可以先把文件从HDFS拉到本地:
在终端执行:

hdfs dfs -get /data/testdata.json ./local_testdata.json

然后用你原本的open()代码读取本地文件:

import json

with open('./local_testdata.json') as data_file:
    json_data = json.load(data_file)

这种方法简单,但不适合大文件(会占用本地磁盘空间)。

3. 通过subprocess调用HDFS命令读取内容

如果不想安装额外库,可以用Python的subprocess模块直接调用HDFS的cat命令获取内容:

import subprocess
import json

try:
    # 执行hdfs cat命令捕获输出
    result = subprocess.run(
        ['hdfs', 'dfs', '-cat', '/data/testdata.json'],
        capture_output=True,
        text=True,
        check=True
    )
    # 解析JSON内容
    json_data = json.loads(result.stdout)
except subprocess.CalledProcessError as e:
    print(f"读取HDFS文件出错: {e.stderr}")

这种方法依赖系统环境里的hdfs命令能正常运行,适合临时脚本,但出错处理会麻烦一些。

内容的提问来源于stack exchange,提问作者S M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:26:30