Pandas基于列值创建新列问题:映射结果全为0排查
问题解决:Pandas映射列全为0的修复方案
错误原因分析
你的代码存在两个关键问题:
- 判断条件与需求不符:你需要把
Never married映射为1,但代码里的判断条件完全没包含这个值,反而错误加入了不需要映射为1的Married-A F spouse present。 - 潜在字符串匹配问题:可能存在字符串前后空格、大小写不一致的情况,导致匹配失败。
修复后的代码方案
基础修复版(修正判断条件)
# 修正判断逻辑,包含需要映射为1的三个类别 df_filtered['is_married'] = df_filtered['marital_status'].apply( lambda status: 1 if status in ["Married-civilian spouse present", "Never married", "Married-spouse absent"] else 0 )
更高效且鲁棒的版本(处理字符串格式问题)
如果担心字符串有空格或大小写问题,可以先清理格式再匹配:
# 先清理字符串:去除首尾空格、统一小写(可选) df_filtered['marital_status_clean'] = df_filtered['marital_status'].str.strip().str.lower() # 定义需要映射为1的目标类别(同步处理格式) target_categories = [ "married-civilian spouse present", "never married", "married-spouse absent" ] # 使用isin方法更高效,避免lambda循环 df_filtered['is_married'] = df_filtered['marital_status_clean'].isin(target_categories).astype(int)
验证步骤
可以先检查marital_status列的实际值,确认匹配是否正确:
# 查看列中所有唯一值 print(df_filtered['marital_status'].unique())
内容的提问来源于stack exchange,提问作者Pranav
相关产品推荐
相关产品推荐

