如何将DataFrame日期字段提取为独立列而非(year, month, day)格式?
解决方法:提取年月日为独立DataFrame列
你的代码里,apply(temp)返回的是每个元素为元组的Series,直接赋值给profile_drop["year","moth","day"]会创建一个多级索引的单列,而不是三个独立列。下面给你几个可行的修正方案:
方案1:修正apply后的赋值逻辑
先保留你的自定义函数,注意把拼写错误的moth改成month,然后把apply结果转为DataFrame再赋值:
def temp(i): i = str(i) year = i[0:4] month = i[4:6] # 修正拼写:moth → month day = i[6:8] return year, month, day # 将apply返回的元组Series转为DataFrame,直接赋值给三个新列 profile_drop[["year", "month", "day"]] = profile_drop["became_member_on"].apply(temp).apply(pd.Series)
方案2:直接用字符串切片(更简洁)
不用写自定义函数,直接对字段做字符串切片生成列:
# 先把字段转成字符串类型(如果原本不是的话) profile_drop["became_member_on"] = profile_drop["became_member_on"].astype(str) # 直接提取年、月、日 profile_drop["year"] = profile_drop["became_member_on"].str[0:4] profile_drop["month"] = profile_drop["became_member_on"].str[4:6] profile_drop["day"] = profile_drop["became_member_on"].str[6:8]
方案3:转成datetime类型提取(推荐)
如果became_member_on是类似20231001的数字/字符串格式,转成datetime类型后提取更规范:
# 转换为datetime格式,指定输入格式为%Y%m%d(年4位+月2位+日2位) profile_drop["became_member_on_dt"] = pd.to_datetime(profile_drop["became_member_on"], format="%Y%m%d") # 提取年、月、日,可根据需求转成字符串或保留整数 profile_drop["year"] = profile_drop["became_member_on_dt"].dt.year.astype(str) profile_drop["month"] = profile_drop["became_member_on_dt"].dt.month.astype(str) profile_drop["day"] = profile_drop["became_member_on_dt"].dt.day.astype(str) # 不需要中间日期列的话可以删除 # profile_drop.drop("became_member_on_dt", axis=1, inplace=True)
内容的提问来源于stack exchange,提问作者yunahwang
相关产品推荐
相关产品推荐

