You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无Pandas时Python心血管病数据分箱与比值计算问题求助

心血管病数据集分析(无法使用Pandas)

任务要求:将height、weight、收缩压(ap_hi)、舒张压(ap_lo)等连续属性划分为5个等宽分箱,计算每个分箱中心血管病患者与非患者的比值,仅保留严格大于1的结果;同时处理cholesterol、glucose、smoking、drinking、activity等分类属性的比值计算。

实现思路

  • 过滤无效数据:移除height、weight、ap_hi、ap_lo字段超出范围的记录。
  • 创建分箱:为连续属性计算分箱,确保最大值边界不遗漏数据点。
  • 计算比值:为每个分箱计算心血管病患者与非患者的比值。
  • 处理分类数据:为cholesterol、glucose等分类属性计算各分类的比值。

完整代码

import csv
from collections import defaultdict

# 按条件过滤数据
def filter_data(data):
    filtered_data = []
    for row in data:
        if not row['height'] or not row['weight'] or not row['ap_hi'] or not row['ap_lo']:
            continue  # 跳过缺失关键值的行
        
        height = int(row['height'])
        weight = float(row['weight'])
        ap_hi = int(row['ap_hi'])
        ap_lo = int(row['ap_lo'])
        
        # 排除超出范围的无效记录
        if height < 150 or height > 200 or weight < 50 or weight > 150 or ap_hi < 80 or ap_hi > 200 or ap_lo < 70 or ap_lo > 140:
            continue
        
        filtered_data.append(row)
    
    return filtered_data

# 为连续属性创建等宽分箱
def create_bins(data, attribute, num_bins=5):
    min_val = min(float(row[attribute]) for row in data)
    max_val = max(float(row[attribute]) for row in data)
    # 扩展最大值边界,确保最大值能被包含进最后一个分箱
    bin_width = (max_val - min_val) / num_bins
    # 最后一个边界设为max_val + 极小值,避免等于max_val的数值无法匹配
    bin_edges = [min_val + i * bin_width for i in range(num_bins)] + [max_val + 1e-9]
    return bin_edges

# 获取数值对应的分箱索引
def get_bin_index(value, bin_edges):
    for i in range(len(bin_edges) - 1):
        if bin_edges[i] <= value < bin_edges[i + 1]:
            return i
    # 兜底返回最后一个分箱索引
    return len(bin_edges) - 2

# 主分析函数
def analyse(gender, age):
    # 加载CSV数据
    data = []
    with open('cardio_train.csv', 'r') as file:
        reader = csv.DictReader(file, delimiter=';')
        for row in reader:
            data.append(row)

    # 过滤无效数据
    filtered_data = filter_data(data)
    
    # 按性别和年龄筛选
    gender_value = 2 if gender == 'F' else 1
    filtered_data = [row for row in filtered_data if int(row['gender']) == gender_value and int(row['age']) // 365 == age]

    if not filtered_data:
        print(f"No data available for {gender.lower()}s aged {age}.")
        return

    # 拆分患病/未患病数据集
    cardio_yes = [row for row in filtered_data if int(row['cardio']) == 1]
    cardio_no = [row for row in filtered_data if int(row['cardio']) == 0]

    # 属性标签映射
    attribute_labels = {
        'height': "Height in category",
        'weight': "Weight in category",
        'ap_hi': "Systolic blood pressure in category",
        'ap_lo': "Diastolic blood pressure in category",
        'cholesterol': "Cholesterol in category",
        'gluc': "Glucose in category"
    }

    # 生成连续属性分箱
    bins = {}
    for attr in ['height', 'weight', 'ap_hi', 'ap_lo']:
        bins[attr] = create_bins(filtered_data, attr)

    # 计算连续属性分箱的比值
    ratios = []
    for attr in ['height', 'weight', 'ap_hi', 'ap_lo']:
        bin_edges = bins[attr]
        for i in range(len(bin_edges) - 1):
            # 统计当前分箱的患病/未患病数量
            yes_count = sum(1 for row in cardio_yes if bin_edges[i] <= float(row[attr]) < bin_edges[i+1])
            no_count = sum(1 for row in cardio_no if bin_edges[i] <= float(row[attr]) < bin_edges[i+1])

            if yes_count + no_count < 5:
                continue  # 跳过样本量过少的分箱

            if no_count > 0:
                ratio = yes_count / no_count
                if ratio > 1:
                    ratios.append((ratio, f"{attribute_labels[attr]} {i + 1} (1 lowest, 5 highest)", attr, i + 1))
            elif yes_count > 0:
                ratios.append((float('inf'), f"{attribute_labels[attr]} {i + 1} (1 lowest, 5 highest)", attr, i + 1))

    # 处理分类属性的比值
    for attr in ['cholesterol', 'gluc', 'smoke', 'alco', 'active']:
        yes_count = defaultdict(int)
        no_count = defaultdict(int)
        for row in cardio_yes:
            yes_count[int(row[attr])] += 1
        for row in cardio_no:
            no_count[int(row[attr])] += 1

        for value in yes_count:
            if value in no_count and no_count[value] > 0:
                ratio = yes_count[value] / no_count[value]
                if ratio > 1:
                    if attr in ['cholesterol', 'gluc']:
                        category = f"{attribute_labels[attr]} {value} (1 lowest, 3 highest)"
                    else:
                        category = (f"{'Being active' if attr == 'active' else 'Smoking' if attr == 'smoke' else 'Drinking'}"
                                    if value == 1 else
                                    f"Not being active" if attr == 'active' else f"Not smoking" if attr == 'smoke' else f"Not drinking")
                    ratios.append((ratio, category, attr, value))
            elif yes_count[value] > 0:
                if attr in ['cholesterol', 'gluc']:
                    category = f"{attribute_labels[attr]} {value} (1 lowest, 3 highest)"
                else:
                    category = (f"{'Being active' if attr == 'active' else 'Smoking' if attr == 'smoke' else 'Drinking'}"
                                if value == 1 else
                                f"Not being active" if attr == 'active' else f"Not smoking" if attr == 'smoke' else f"Not drinking")
                ratios.append((float('inf'), category, attr, value))

    # 按比值降序、属性优先级、分箱序号排序
    all_attributes = ['height', 'weight', 'ap_hi', 'ap_lo', 'cholesterol', 'gluc', 'smoke', 'alco', 'active']
    ratios.sort(key=lambda x: (-x[0], all_attributes.index(x[2]), x[3]))

    # 输出结果
    if ratios:
        print(f"The following might particularly contribute to cardio problems for {'females' if gender == 'F' else 'males'} aged {age}:")
        for ratio, description, _, _ in ratios:
            ratio_str = 'inf' if ratio == float('inf') else f"{ratio:.2f}"
            print(f"   {ratio_str}: {description}")

# 示例调用
analyse('F', 43)

遇到的问题

  • 分箱分类错误:部分ap_hi、ap_lo等值未被正确归类,例如应属于第5分箱的值被分到第3分箱。
  • 比值结果不符:计算的比值与预期结果存在差异,例如舒张压第4分箱预期比值为17.88,实际得到11.25。
  • 最大值处理问题:即使设置了max_adjustment=0.1,仍存在高值分箱归类错误的情况。

求助问题

  • 如何确保连续属性分箱能准确覆盖数据范围,尤其是极值?
  • 如何调整比值计算逻辑以匹配预期输出?
  • 有没有更优的方法处理边界案例,确保最大值被正确归类?

内容的提问来源于stack exchange,提问作者DhirajInAu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 01:54:58