You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark读取CSV时所有列显示为String,如何获取真实数据类型?

PySpark读取CSV时如何获取真实数据类型?

你目前用默认方式读取CSV文件时,Spark会将所有列默认识别为String类型,而且第一行的表头也没被当作列名,导致列名显示为_c0、_c1这类自动命名的标识。要获取各列的真实数据类型,可以通过以下两种方式解决:

1. 启用自动类型推断与表头识别

在读取CSV时添加header=True和inferSchema=True参数,让Spark自动识别表头,并推断每列的真实数据类型:

df_pyspark = spark.read.csv("sample_data.csv", header=True, inferSchema=True)
# 查看各列的数据类型
df_pyspark.printSchema()

执行后,printSchema()会输出各列的真实类型,比如id会被识别为整数类型,phone会被识别为长整型,示例输出如下:

root
 |-- id: integer (nullable = true)
 |-- first_name: string (nullable = true)
 |-- last_name: string (nullable = true)
 |-- email: string (nullable = true)
 |-- gender: string (nullable = true)
 |-- phone: long (nullable = true)

2. 手动指定Schema(更精准控制)

如果自动推断的类型不符合预期(比如phone需要特定数值类型,或者某些列有特殊格式),可以手动定义Schema来强制指定各列类型:

from pyspark.sql.types import StructType, StructField, IntegerType, StringType, LongType

# 定义自定义Schema
custom_schema = StructType([
    StructField("id", IntegerType(), nullable=True),
    StructField("first_name", StringType(), nullable=True),
    StructField("last_name", StringType(), nullable=True),
    StructField("email", StringType(), nullable=True),
    StructField("gender", StringType(), nullable=True),
    StructField("phone", LongType(), nullable=True)
])

# 使用自定义Schema读取CSV
df_pyspark = spark.read.csv("sample_data.csv", header=True, schema=custom_schema)
df_pyspark.printSchema()

这种方式适合数据格式固定的场景,能确保数据类型完全符合需求。

内容的提问来源于stack exchange,提问作者Darshan patil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 06:25:18