Spark读取混合类型Parquet写入时year列类型转换错误求助
Your explicit cast isn't working because you're using the IntegerType class instead of an instance of it. In PySpark, you need to pass IntegerType() (with parentheses) or the string "integer" to the cast() method.
But even with correct casting, if the INT64 files contain values larger than the 32-bit integer limit (-2^31 to 2^31-1), you'll hit overflow issues or nulls. A more reliable fix is to use LongType (Spark's 64-bit integer type) for the year column—it can handle values from both int and INT64 files without conversion problems.
Fix 1: Correct the cast syntax (if all values fit in 32-bit integers)
Update your cast code to use valid syntax:
from pyspark.sql.types import IntegerType from pyspark.sql.functions import col df = df.select(col('year').cast(IntegerType())) # Or use the string shorthand: # df = df.select(col('year').cast("integer")) df.write.parquet('gs://myresults/')
Fix 2: Switch to LongType to avoid type mismatches entirely
Modify your schema to use LongType for the year column when reading—this matches the INT64 type from the 20% of files:
from pyspark.sql.types import LongType, StructType, StructField # Update your schema definition myschema = StructType([ # Add your other fields here StructField("year", LongType(), nullable=True) ]) # Read the files with the updated schema df = spark.read.schema(myschema).parquet('gs://path/myfiles/') # Write directly without casting df.write.parquet('gs://myresults/')
This works because LongType can accommodate all values from both 32-bit int and 64-bit INT64, eliminating the type mismatch that causes the write error.
If you absolutely need the output column to be IntegerType, first verify no values exceed the 32-bit range with this check:
from pyspark.sql.functions import col # Show rows where year is outside 32-bit integer bounds df.filter((col("year") < -2147483648) | (col("year") > 2147483647)).show()
If there are no such rows, use Fix 1. If there are, you'll need to either adjust those values or stick with LongType to preserve data integrity.
内容的提问来源于stack exchange,提问作者tsh

