Python/pandas统计数据集满足测试判定条件的元素数量及分类占比
实现步骤和代码
1. 导入依赖并构造数据集
import pandas as pd df = pd.DataFrame({'ID': {0: 'GF342', 1: 'IF874', 2: 'FH386', 3: 'KJ190', 4: 'TY748', 5: 'YT947', 6: 'DF063', 7: 'ET512', 8: 'GC714', 9: 'SD978', 10: 'EF472', 11: 'PL489', 12: 'AZ315', 13: 'OL821', 14: 'HN765', 15: 'ED589'}, 'Location': {0: 'Q1', 1: 'Q3', 2: 'Q1', 3: 'Q3', 4: 'Q3', 5: 'Q4', 6: 'Q3', 7: 'Q1', 8: 'Q2', 9: 'Q3', 10: 'Q1', 11: 'Q2', 12: 'Q1', 13: 'Q1', 14: 'Q3', 15: 'Q1'}, 'NEW': {0: 'YES', 1: 'NO', 2: 'NO', 3: 'YES', 4: 'YES', 5: 'NO', 6: 'NO', 7: 'YES', 8: 'NO', 9: 'NO', 10: 'NO', 11: 'YES', 12: 'NO', 13: 'YES', 14: 'YES', 15: 'YES'}, 'YEAR': {0: 2021, 1: 2018, 2: 2019, 3: 2021, 4: 2021, 5: 2019, 6: 2019, 7: 2021, 8: 2018, 9: 2019, 10: 2018, 11: 2021, 12: 2018, 13: 2021, 14: 2021, 15: 2021}, 'PT1': {0: '', 1: 'NOT_TESTED', 2: '', 3: 'NOT_FINISHED', 4: '', 5: '', 6: '180', 7: '', 8: '', 9: '', 10: '', 11: '', 12: 'TOO_LOW', 13: '', 14: '155', 15: ''}, 'PT2': {0: '', 1: '', 2: '', 3: '', 4: '', 5: 'TOO_LOW', 6: '', 7: '', 8: '160', 9: 'TOO_LOW', 10: '', 11: '', 12: '', 13: '', 14: '', 15: ''}, 'PT3': {0: '', 1: 'TOO_LOW', 2: '', 3: 'TOO_LOW', 4: '', 5: '', 6: '', 7: '', 8: '', 9: '', 10: '', 11: 'NOT_FINISHED', 12: '', 13: '185', 14: '', 15: '165'}, 'PT4': {0: '', 1: '', 2: '', 3: '', 4: '', 5: 165.0, 6: '', 7: '', 8: '', 9: '', 10: '', 11: '', 12: 180.0, 13: '', 14: '', 15: ''}})
2. 标记每行是否为已测试
先提取所有PT列,然后按规则逐行判断:
pt_cols = ['PT1', 'PT2', 'PT3', 'PT4'] def check_is_tested(row): for col in pt_cols: val = row[col] # 跳过空值 if pd.isna(val) or str(val).strip() == '': continue # 判断是否为too_low if str(val).strip().upper() == 'TOO_LOW': return True # 判断是否为150-190区间的数值 try: num_val = float(val) if 150 <= num_val <= 190: return True except ValueError: continue return False # 新增布尔列标记是否已测试 df['is_tested'] = df.apply(check_is_tested, axis=1)
3. 计算统计指标
# 总样本量 total_cnt = len(df) # 已测试总数量及占比 tested_cnt = df['is_tested'].sum() tested_ratio = round(tested_cnt / total_cnt * 100, 2) print(f"已测试元素总数量:{tested_cnt},占全部元素比例:{tested_ratio}%") # 已测试样本切片 tested_df = df[df['is_tested']] # 新测试元素数量及占比 new_tested_cnt = (tested_df['NEW'] == 'YES').sum() new_tested_ratio = round(new_tested_cnt / tested_cnt * 100, 2) print(f"已测试中新测试元素数量:{new_tested_cnt},占已测试元素比例:{new_tested_ratio}%") # 旧测试元素数量及占比 old_tested_cnt = (tested_df['NEW'] == 'NO').sum() old_tested_ratio = round(old_tested_cnt / tested_cnt * 100, 2) print(f"已测试中旧测试元素数量:{old_tested_cnt},占已测试元素比例:{old_tested_ratio}%")
运行结果(示例数据集)
已测试元素总数量:10,占全部元素比例:62.5% 已测试中新测试元素数量:4,占已测试元素比例:40.0% 已测试中旧测试元素数量:6,占已测试元素比例:60.0%
内容的提问来源于stack exchange,提问作者Boston Co
相关产品推荐
相关产品推荐

