You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyArrow导出CSV遇浮点精度问题,求类似Pandas的float_format方案

解决PyArrow导出CSV的浮点精度显示问题

核心方案:用CsvWriteOptions配置浮点格式

PyArrow的csv.write_csv支持通过CsvWriteOptions设置浮点数输出格式,完全对应Pandas里float_format='%g'的作用,还能保留PyArrow的高速导出优势。

直接写代码实现:

import pyarrow.csv as csv
import pyarrow as pa

# 假设已经构造好要导出的pa.Table对象table
write_options = csv.CsvWriteOptions(float_format='%g')
csv.write_csv(table, "animals.csv", write_options=write_options)

进阶:特定列单独格式化

如果需要给不同浮点列设置不同格式,可以先通过PyArrow计算API把浮点列转成格式化后的字符串列再导出(注意:此方法会将列类型转为字符串,适合下游不需要数值类型的场景):

# 示例:给名为"price"的浮点列设置保留两位小数的格式
target_col_idx = table.column_names.index("price")
formatted_col = pa.compute.format_float(table["price"], format='%.2f')
table = table.set_column(target_col_idx, "price", formatted_col)

# 导出处理后的表
csv.write_csv(table, "animals.csv")

补充说明

  • 第一种方案是最优解:既不损失PyArrow的导出速度,又能让浮点数输出和Pandas用%g格式化的效果一致,避免出现1.12999999999这类冗余精度的显示。
  • %g格式会自动去掉浮点数末尾的无效零,还能在科学计数法和普通十进制表示间自动切换,兼顾可读性和精度。

内容的提问来源于stack exchange,提问作者dogfood

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 00:50:37