You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark读取CSV时Double类型数值自动舍入问题

问题原因

Double类型是64位浮点数,仅能精确表示二进制有效位≤53位的整数(对应十进制约15-17位)。9159613164437314这个数的二进制有效位超过了53位,无法被Double精确存储,会被自动舍入到最近的可精确表示值——也就是9159613164437312。

你提到JSON中的值显示为9159613164437314.0,大概率是因为JSON读取时,Spark将其解析为字符串或DecimalType而非Double,从而保留了精度。

解决方案

1. 手动指定Schema,用DecimalType读取CSV

直接在读取CSV时将目标列定义为DecimalType(指定足够大的精度),跳过Double类型的转换:

from pyspark.sql.types import StructType, StructField, DecimalType

# 定义Schema,将目标列设为DecimalType(38, 0),支持最大38位整数
csv_schema = StructType([
    StructField("target_column", DecimalType(38, 0), nullable=True),
    # 其他列根据实际情况添加
])

# 读取CSV文件
df = spark.read.csv("your_file.csv", schema=csv_schema, header=True)

2. 先以StringType读取,再转DecimalType

如果不确定数值范围,先按字符串读取避免精度丢失,再转换为DecimalType:

from pyspark.sql.types import StructType, StructField, StringType, DecimalType

csv_schema = StructType([
    StructField("target_column", StringType(), nullable=True),
    # 其他列...
])

df = spark.read.csv("your_file.csv", schema=csv_schema, header=True)

# 转换为高精度Decimal类型
df = df.withColumn("target_column", df["target_column"].cast(DecimalType(38, 0)))

3. 关闭自动类型推断(可选)

如果依赖Spark自动推断类型,它会默认将大数值推断为Double,可关闭推断强制手动指定Schema:

df = spark.read.csv(
    "your_file.csv",
    header=True,
    inferSchema=False,  # 关闭自动推断
    schema=csv_schema  # 手动传入Schema
)

内容的提问来源于stack exchange,提问作者Xi12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 10:16:11