无Pandas时Python心血管病数据分箱与比值计算问题求助
心血管病数据集分析(无法使用Pandas)
任务要求:将height、weight、收缩压(ap_hi)、舒张压(ap_lo)等连续属性划分为5个等宽分箱,计算每个分箱中心血管病患者与非患者的比值,仅保留严格大于1的结果;同时处理cholesterol、glucose、smoking、drinking、activity等分类属性的比值计算。
实现思路
- 过滤无效数据:移除height、weight、ap_hi、ap_lo字段超出范围的记录。
- 创建分箱:为连续属性计算分箱,确保最大值边界不遗漏数据点。
- 计算比值:为每个分箱计算心血管病患者与非患者的比值。
- 处理分类数据:为cholesterol、glucose等分类属性计算各分类的比值。
完整代码
import csv from collections import defaultdict # 按条件过滤数据 def filter_data(data): filtered_data = [] for row in data: if not row['height'] or not row['weight'] or not row['ap_hi'] or not row['ap_lo']: continue # 跳过缺失关键值的行 height = int(row['height']) weight = float(row['weight']) ap_hi = int(row['ap_hi']) ap_lo = int(row['ap_lo']) # 排除超出范围的无效记录 if height < 150 or height > 200 or weight < 50 or weight > 150 or ap_hi < 80 or ap_hi > 200 or ap_lo < 70 or ap_lo > 140: continue filtered_data.append(row) return filtered_data # 为连续属性创建等宽分箱 def create_bins(data, attribute, num_bins=5): min_val = min(float(row[attribute]) for row in data) max_val = max(float(row[attribute]) for row in data) # 扩展最大值边界,确保最大值能被包含进最后一个分箱 bin_width = (max_val - min_val) / num_bins # 最后一个边界设为max_val + 极小值,避免等于max_val的数值无法匹配 bin_edges = [min_val + i * bin_width for i in range(num_bins)] + [max_val + 1e-9] return bin_edges # 获取数值对应的分箱索引 def get_bin_index(value, bin_edges): for i in range(len(bin_edges) - 1): if bin_edges[i] <= value < bin_edges[i + 1]: return i # 兜底返回最后一个分箱索引 return len(bin_edges) - 2 # 主分析函数 def analyse(gender, age): # 加载CSV数据 data = [] with open('cardio_train.csv', 'r') as file: reader = csv.DictReader(file, delimiter=';') for row in reader: data.append(row) # 过滤无效数据 filtered_data = filter_data(data) # 按性别和年龄筛选 gender_value = 2 if gender == 'F' else 1 filtered_data = [row for row in filtered_data if int(row['gender']) == gender_value and int(row['age']) // 365 == age] if not filtered_data: print(f"No data available for {gender.lower()}s aged {age}.") return # 拆分患病/未患病数据集 cardio_yes = [row for row in filtered_data if int(row['cardio']) == 1] cardio_no = [row for row in filtered_data if int(row['cardio']) == 0] # 属性标签映射 attribute_labels = { 'height': "Height in category", 'weight': "Weight in category", 'ap_hi': "Systolic blood pressure in category", 'ap_lo': "Diastolic blood pressure in category", 'cholesterol': "Cholesterol in category", 'gluc': "Glucose in category" } # 生成连续属性分箱 bins = {} for attr in ['height', 'weight', 'ap_hi', 'ap_lo']: bins[attr] = create_bins(filtered_data, attr) # 计算连续属性分箱的比值 ratios = [] for attr in ['height', 'weight', 'ap_hi', 'ap_lo']: bin_edges = bins[attr] for i in range(len(bin_edges) - 1): # 统计当前分箱的患病/未患病数量 yes_count = sum(1 for row in cardio_yes if bin_edges[i] <= float(row[attr]) < bin_edges[i+1]) no_count = sum(1 for row in cardio_no if bin_edges[i] <= float(row[attr]) < bin_edges[i+1]) if yes_count + no_count < 5: continue # 跳过样本量过少的分箱 if no_count > 0: ratio = yes_count / no_count if ratio > 1: ratios.append((ratio, f"{attribute_labels[attr]} {i + 1} (1 lowest, 5 highest)", attr, i + 1)) elif yes_count > 0: ratios.append((float('inf'), f"{attribute_labels[attr]} {i + 1} (1 lowest, 5 highest)", attr, i + 1)) # 处理分类属性的比值 for attr in ['cholesterol', 'gluc', 'smoke', 'alco', 'active']: yes_count = defaultdict(int) no_count = defaultdict(int) for row in cardio_yes: yes_count[int(row[attr])] += 1 for row in cardio_no: no_count[int(row[attr])] += 1 for value in yes_count: if value in no_count and no_count[value] > 0: ratio = yes_count[value] / no_count[value] if ratio > 1: if attr in ['cholesterol', 'gluc']: category = f"{attribute_labels[attr]} {value} (1 lowest, 3 highest)" else: category = (f"{'Being active' if attr == 'active' else 'Smoking' if attr == 'smoke' else 'Drinking'}" if value == 1 else f"Not being active" if attr == 'active' else f"Not smoking" if attr == 'smoke' else f"Not drinking") ratios.append((ratio, category, attr, value)) elif yes_count[value] > 0: if attr in ['cholesterol', 'gluc']: category = f"{attribute_labels[attr]} {value} (1 lowest, 3 highest)" else: category = (f"{'Being active' if attr == 'active' else 'Smoking' if attr == 'smoke' else 'Drinking'}" if value == 1 else f"Not being active" if attr == 'active' else f"Not smoking" if attr == 'smoke' else f"Not drinking") ratios.append((float('inf'), category, attr, value)) # 按比值降序、属性优先级、分箱序号排序 all_attributes = ['height', 'weight', 'ap_hi', 'ap_lo', 'cholesterol', 'gluc', 'smoke', 'alco', 'active'] ratios.sort(key=lambda x: (-x[0], all_attributes.index(x[2]), x[3])) # 输出结果 if ratios: print(f"The following might particularly contribute to cardio problems for {'females' if gender == 'F' else 'males'} aged {age}:") for ratio, description, _, _ in ratios: ratio_str = 'inf' if ratio == float('inf') else f"{ratio:.2f}" print(f" {ratio_str}: {description}") # 示例调用 analyse('F', 43)
遇到的问题
- 分箱分类错误:部分ap_hi、ap_lo等值未被正确归类,例如应属于第5分箱的值被分到第3分箱。
- 比值结果不符:计算的比值与预期结果存在差异,例如舒张压第4分箱预期比值为17.88,实际得到11.25。
- 最大值处理问题:即使设置了
max_adjustment=0.1,仍存在高值分箱归类错误的情况。
求助问题
- 如何确保连续属性分箱能准确覆盖数据范围,尤其是极值?
- 如何调整比值计算逻辑以匹配预期输出?
- 有没有更优的方法处理边界案例,确保最大值被正确归类?
内容的提问来源于stack exchange,提问作者DhirajInAu
相关产品推荐
相关产品推荐

