如何为Example.Data数据集添加符合指定规则的新列?
解决方案
以下是主流数据处理工具的实现方法,涵盖R和Python(Pandas)两种常用场景:
使用R语言
利用cut()函数可以快速实现区间分箱,匹配你需要的分类规则:
# 假设数据集已加载为Example.Data Example.Data$Education_Level <- cut( Example.Data$Educ, breaks = c(12, 16, 17, 19, Inf), labels = c("HighSchool", "College", "Masters", "Doctorate"), include.lowest = TRUE, # 包含12这个下限值 right = FALSE # 设置区间为左闭右开,匹配规则中的范围 )
使用Python(Pandas)
借助Pandas的pd.cut()函数完成分箱操作:
import pandas as pd # 假设数据集已加载为Example_Data bins = [12, 16, 17, 19, float('inf')] labels = ["HighSchool", "College", "Masters", "Doctorate"] Example_Data['Education_Level'] = pd.cut( Example_Data['Educ'], bins=bins, labels=labels, include_lowest=True, # 确保12被包含在第一个区间 right=False # 左闭右开区间,符合规则定义 )
备注:你提供的现有数据集示例中未包含
Educ字段,实际操作前请确保数据集存在该数值型字段,否则上述代码无法正常运行。
内容的提问来源于stack exchange,提问作者ellen
相关产品推荐
相关产品推荐

