如何筛选出两次测量中均存在的树木观测数据?
树木生存状态分类解决方案
问题分析
需要基于State、County、Plot、Tree这组唯一标识,区分树木在t1到t2期间的三种状态:
- 跨t1/t2存在:存活(标记1)
- 仅t1存在:死亡(标记0)
- 仅t2存在:新生长树木(目标数据集无需保留)
实现方案(Python Pandas)
1. 数据准备
先将原始数据加载到DataFrame:
import pandas as pd # 原始数据集 data = { 'State': [1,1,1,1,1,1,1], 'County': [9,9,9,9,9,9,9], 'Plot': [1,1,1,1,1,1,1], 'Tree': [1,2,3,1,2,4,5], 'Meas_yr': ['t1','t1','t1','t2','t2','t2','t2'] } df = pd.DataFrame(data)
2. 标记树木生存状态
按唯一标识分组,统计每个树木的观测年份数量,再判断状态:
# 按唯一标识分组,获取每个树木的所有观测年份 tree_year_groups = df.groupby(['State', 'County', 'Plot', 'Tree'])['Meas_yr'].unique() # 生成状态映射:同时有t1和t2则为1,只有t1则为0,只有t2则标记为'new' status_map = {} for tree_id, years in tree_year_groups.items(): if 't1' in years and 't2' in years: status_map[tree_id] = 1 elif 't1' in years: status_map[tree_id] = 0 else: status_map[tree_id] = 'new' # 将状态映射合并回原数据集 df['tree_survival'] = df.apply(lambda row: status_map[(row['State'], row['County'], row['Plot'], row['Tree'])], axis=1)
3. 生成目标数据集
过滤掉仅t2存在的新树木记录:
# 筛选出非'new'的记录 target_df = df[df['tree_survival'] != 'new'].sort_values(by=['Meas_yr', 'Tree']) # 重置索引 target_df = target_df.reset_index(drop=True)
4. 查看结果
输出目标数据集:
print(target_df)
输出结果与需求一致:
State County Plot Tree Meas_yr tree_survival 0 1 9 1 1 t1 1 1 1 9 1 2 t1 1 2 1 9 1 3 t1 0 3 1 9 1 1 t2 1 4 1 9 1 2 t2 1
关键逻辑说明
- 利用分组操作锁定每个树木的时间跨度,确保状态判断的准确性
- 通过映射字典快速给每条记录标记生存状态,避免循环遍历的低效
- 最后过滤新树木,完全匹配需求输出结构
内容的提问来源于stack exchange,提问作者sakar299
相关产品推荐
相关产品推荐

