如何在PySpark与Python中将年份列转换为指定格式?
将年份转换为YYYY-YY格式的实现方案
要实现将单个年份(如2022、2021)转换为YYYY-YY格式(2022→2022-23,2021→2021-22),以下是PySpark和Python环境下的具体实现方案:
PySpark 实现
假设数据列名为year,支持整数或字符串类型的年份输入,核心逻辑是计算年份+1后的最后两位,再与原年份拼接:
from pyspark.sql import SparkSession from pyspark.sql.functions import col, format_string # 初始化Spark会话 spark = SparkSession.builder.appName("YearFormatConversion").getOrCreate() # 创建示例DataFrame(混合整数、字符串类型的年份) sample_data = [(2022,), (2021,), ("2020",), (2099,)] df = spark.createDataFrame(sample_data, ["year"]) # 转换逻辑:统一转为整数→生成格式化列 formatted_df = df.withColumn("year_int", col("year").cast("int")) \ .withColumn( "formatted_year", format_string("%d-%02d", col("year_int"), (col("year_int") + 1) % 100) ) \ .drop("year_int") # 查看结果 formatted_df.show()
说明
col("year").cast("int"):统一将年份转为整数类型,避免字符串处理的异常(year_int + 1) % 100:确保年份+1后取最后两位,比如2099会得到00%02d:强制后两位为两位数格式,避免出现2000-1这类不规范输出
Python 实现
分纯Python处理单值/批量数据,以及Pandas处理表格数据两种场景:
1. 纯Python 处理单值或批量列表
定义通用函数支持整数、字符串类型的年份输入:
def format_fiscal_year(year): year_int = int(year) next_year_suffix = str(year_int + 1)[-2:] return f"{year_int}-{next_year_suffix}" # 测试单值 print(format_fiscal_year(2022)) # 输出: 2022-23 print(format_fiscal_year("2021")) # 输出: 2021-22 print(format_fiscal_year(2099)) # 输出: 2099-00 # 批量处理列表 raw_years = [2022, "2021", 2020, 2099] formatted_years = [format_fiscal_year(y) for y in raw_years] print(formatted_years) # 输出: ['2022-23', '2021-22', '2020-21', '2099-00']
2. Pandas 处理DataFrame列
适合处理表格数据,优先使用矢量化操作提升效率:
import pandas as pd # 创建示例DataFrame df = pd.DataFrame({"year": [2022, 2021, "2020", 2099]}) # 统一转为整数类型 df["year"] = df["year"].astype(int) # 方式1:使用apply(适合简单逻辑) df["formatted_year"] = df["year"].apply(lambda x: f"{x}-{str(x+1)[-2:]}") # 方式2:矢量化操作(大数据量下效率更高) df["formatted_year"] = df["year"].astype(str) + "-" + (df["year"] + 1).astype(str).str[-2:] # 查看结果 print(df)
内容的提问来源于stack exchange,提问作者user19814628
相关产品推荐
相关产品推荐

