pandas指定列截取子串、Date of Birth列求最值及转分类变量咨询
问题解决方法
报错原因说明
Date of Birth列被pandas解析为datetime64类型而非字符串类型,因此不存在.str访问器属性- 第二种写法错误引用了不存在的
Name列,同时逐行遍历修改pandas数据属于极低效的操作,不推荐使用
步骤1:统一日期列格式
首先将出生日期列转换为标准的pandas日期格式:
import pandas as pd # errors='coerce'会把不符合格式的内容转为空值,避免报错 dataset1['Date of Birth'] = pd.to_datetime(dataset1['Date of Birth'], errors='coerce')
步骤2:获取日期最大、最小值
直接调用Series内置的统计方法即可:
# 完整日期的最大最小值 min_dob = dataset1['Date of Birth'].min() max_dob = dataset1['Date of Birth'].max() # 若只需年份的最大最小值 min_birth_year = dataset1['Date of Birth'].dt.year.min() max_birth_year = dataset1['Date of Birth'].dt.year.max()
步骤3:按年份区间分组生成分类变量
使用pandas的cut方法快速完成分箱分组,无需自行遍历:
# 先提取出生年份列 dataset1['birth_year'] = dataset1['Date of Birth'].dt.year # 自定义分箱边界,可根据实际年份范围调整 bins = [1969, 1974, 1980, 1985, 1990, 1995, 2000] # 自定义分组标签,和上方分箱区间一一对应 labels = [0, 1, 2, 3, 4, 5] # 生成分组列,include_lowest=True保证左边界数值被正确归入对应区间 dataset1['year_group'] = pd.cut( dataset1['birth_year'], bins=bins, labels=labels, include_lowest=True )
如果你的Date of Birth列确实需要用字符串方式截取年份,可先强制转为字符串类型再操作:
dataset1['birth_year'] = dataset1['Date of Birth'].astype(str).str[:4].astype(int)
内容的提问来源于stack exchange,提问作者PassaroBagante
相关产品推荐
相关产品推荐

