评估调查流失概率与现金补助干预的关联
问题
我有一组非平衡面板调查数据,每行记录个体、访谈日期及当期劳动力市场状态——部分个体因停止应答,导致不同个体的观测次数存在差异。数据涵盖部分个体在2019-09-03(干预日期)被随机给予现金补助前后的信息,想检验:获得现金补助后,个体是否更大概率停止应答调查?也就是分析调查流失概率与date变量的关联,不确定具体实现方法。
数据示例:
individual date labor_status cash_benefit Kenny 2018-09-02 unemployed 0 Kenny 2019-09-03 unemployed 1 Kenny 2020-09-07 employed 1 Kenny 2021-09-13 employed 1 Cartman 2018-09-03 unemployed 0 Cartman 2019-09-06 unemployed 1 Cartman 2020-09-08 N/A 1 Cartman 2021-09-08 N/A 1 Mackey 2018-09-03 employed 0 Mackey 2019-09-04 unemployed 0 Mackey 2020-09-08 employed 0 Mackey 2021-09-13 employed 0
分析思路与实现步骤
1. 定义「流失」变量
首先明确“停止应答”的判定规则:个体最后一次给出有效劳动力状态(labor_status不为N/A)后,后续访谈全为N/A,或无后续观测。按以下步骤构造流失标识:
- 对每个个体,找出其最后一次有效应答的日期(即
labor_status != 'N/A'的最大日期) - 若个体在干预日期(2019-09-03)之后的观测中出现
N/A,且后续无有效应答,则标记attrition=1,否则为0
2. 构造核心解释变量
区分处理组(获现金补助)和对照组,以及干预前后的时间窗口:
treatment:二进制变量,cash_benefit=1的个体为1,否则为0(随机分配,直接用该变量标识处理组)post_treatment:二进制变量,访谈日期在2019-09-03及之后的为1,否则为0- 交互项
treatment*post_treatment:核心关注变量,衡量处理组在干预后的流失概率变化
3. 模型选择与代码实现
针对面板数据的流失分析,常用两种方法,可按需选择:
方法一:线性概率模型(LPM)+ 个体固定效应
适合初步分析,结果直观易解释,能控制个体层面固定异质性:
import pandas as pd import linearmodels as plm # 数据预处理:转换日期格式 df['date'] = pd.to_datetime(df['date']) # 标记干预后时间段 df['post_treatment'] = (df['date'] >= '2019-09-03').astype(int) # 构造核心交互项 df['treatment_post'] = df['cash_benefit'] * df['post_treatment'] # 标记有效应答 df['valid_response'] = (df['labor_status'] != 'N/A').astype(int) # 找出每个个体最后一次有效应答的日期 last_valid_date = df[df['valid_response'] == 1].groupby('individual')['date'].max().reset_index() last_valid_date.columns = ['individual', 'last_valid'] df = df.merge(last_valid_date, on='individual') # 构造流失变量:当前观测在最后一次有效应答之后,且为无效应答 df['attrition'] = ((df['date'] > df['last_valid']) & (df['valid_response'] == 0)).astype(int) # 拟合个体固定效应模型 model = plm.PanelOLS.from_formula( 'attrition ~ treatment_post + post_treatment + cash_benefit + EntityEffects', data=df.set_index(['individual', 'date']) ) # 按个体聚类计算标准误 results = model.fit(cov_type='clustered', cluster_entity=True) print(results.summary())
方法二:Cox比例风险模型(生存分析)
若想分析「个体何时发生流失」而非单纯流失概率,生存分析更合适。将个体从首次访谈至流失的时间作为生存时间,未流失个体视为删失数据:
import lifelines # 计算每个个体首次访谈日期 first_date = df.groupby('individual')['date'].min().reset_index() first_date.columns = ['individual', 'first_date'] df = df.merge(first_date, on='individual') # 计算从首次访谈至今的月份数(作为生存时间单位) df['time_since_first'] = (df['date'] - df['first_date']).dt.days // 30 # 找出每个个体首次出现无效应答的时间 attrition_time = df[df['valid_response'] == 0].groupby('individual')['time_since_first'].min().reset_index() attrition_time.columns = ['individual', 'attrition_time'] # 找出每个个体最后一次访谈的时间 last_time = df.groupby('individual')['time_since_first'].max().reset_index() last_time.columns = ['individual', 'last_time'] # 构造生存数据集 survival_data = pd.merge(last_time, attrition_time, on='individual', how='left') # 标记是否发生流失事件 survival_data['event'] = survival_data['attrition_time'].notna().astype(int) # 生存时间:发生流失则取首次流失时间,未流失则取最后一次访谈时间 survival_data['T'] = survival_data.apply(lambda x: x['attrition_time'] if x['event'] else x['last_time'], axis=1) # 合并处理组标识 survival_data = survival_data.merge(df[['individual', 'cash_benefit']].drop_duplicates(), on='individual') # 拟合Cox比例风险模型 cox_model = lifelines.CoxPHFitter() cox_model.fit(survival_data, duration_col='T', event_col='event', formula='cash_benefit') cox_model.print_summary()
4. 结果解读
- 线性模型中:若
treatment_post的系数显著为正,说明处理组在干预后的流失概率显著高于对照组 - Cox模型中:若
cash_benefit的风险比(HR)大于1且显著,说明接受现金补助的个体流失风险更高
内容的提问来源于stack exchange,提问作者nesta1990
相关产品推荐
相关产品推荐

