You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark读取混合类型Parquet写入时year列类型转换错误求助

Your explicit cast isn't working because you're using the IntegerType class instead of an instance of it. In PySpark, you need to pass IntegerType() (with parentheses) or the string "integer" to the cast() method.

But even with correct casting, if the INT64 files contain values larger than the 32-bit integer limit (-2^31 to 2^31-1), you'll hit overflow issues or nulls. A more reliable fix is to use LongType (Spark's 64-bit integer type) for the year column—it can handle values from both int and INT64 files without conversion problems.

Fix 1: Correct the cast syntax (if all values fit in 32-bit integers)

Update your cast code to use valid syntax:

from pyspark.sql.types import IntegerType
from pyspark.sql.functions import col

df = df.select(col('year').cast(IntegerType()))
# Or use the string shorthand:
# df = df.select(col('year').cast("integer"))

df.write.parquet('gs://myresults/')

Fix 2: Switch to LongType to avoid type mismatches entirely

Modify your schema to use LongType for the year column when reading—this matches the INT64 type from the 20% of files:

from pyspark.sql.types import LongType, StructType, StructField

# Update your schema definition
myschema = StructType([
    # Add your other fields here
    StructField("year", LongType(), nullable=True)
])

# Read the files with the updated schema
df = spark.read.schema(myschema).parquet('gs://path/myfiles/')

# Write directly without casting
df.write.parquet('gs://myresults/')

This works because LongType can accommodate all values from both 32-bit int and 64-bit INT64, eliminating the type mismatch that causes the write error.

If you absolutely need the output column to be IntegerType, first verify no values exceed the 32-bit range with this check:

from pyspark.sql.functions import col

# Show rows where year is outside 32-bit integer bounds
df.filter((col("year") < -2147483648) | (col("year") > 2147483647)).show()

If there are no such rows, use Fix 1. If there are, you'll need to either adjust those values or stick with LongType to preserve data integrity.

内容的提问来源于stack exchange,提问作者tsh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 16:48:10