如何用Python将Feather文件转为CSV?读取时遇ArrowInvalid错误
Feather转CSV:解决读取错误与转换方法
一、排查"File is too small to be a well-formed file"错误
- 核对文件路径:确认
coffee_ratings_full.feather的路径拼写正确,文件确实存在于指定目录 - 验证文件完整性:若文件为下载所得,重新下载排查损坏问题;若文件大小仅几字节,说明文件已损坏或未正确生成
- 确认文件格式:检查文件是否真的是Feather格式——部分文件可能被错误修改后缀名,实际并非Feather文件
二、正确的Feather转CSV流程
不管使用feather、pyarrow还是pandas,核心逻辑都是读取Feather数据后导出为CSV,你之前觉得“仅能读取无法转换”,大概率是缺少了保存为CSV的步骤,以下是具体实现:
方案1:使用Pandas(最简便)
import pandas as pd # 读取Feather文件(需确保文件无损坏) df = pd.read_feather("coffee_ratings_full.feather") # 导出为CSV,index=False避免写入冗余索引列 df.to_csv("coffee_ratings_full.csv", index=False)
方案2:使用PyArrow
import pyarrow.feather as feather import pyarrow.csv as csv # 读取Feather数据为Arrow Table table = feather.read_table("coffee_ratings_full.feather") # 将Table写入CSV文件 csv.write_csv(table, "coffee_ratings_full.csv")
三、文件有效性验证(读取失败时用)
若确认文件路径和完整性无问题,但仍无法读取,可用PyArrow底层API验证文件是否为标准Feather/Arrow格式:
import pyarrow as pa with pa.memory_map("coffee_ratings_full.feather", "r") as source: try: reader = pa.ipc.open_file(source) print("文件为有效Arrow/Feather格式") print(f"数据批次数量:{reader.num_record_batches()}") except Exception as e: print(f"文件无效:{str(e)}")
内容的提问来源于stack exchange,提问作者Zeina Yousri
相关产品推荐
相关产品推荐

