如何将带表头、体重分类的字符串列表转换为规范pandas DataFrame
解决方案
核心思路是先单独处理表头拼接,再批量清洗业务数据后生成DataFrame,完整可运行代码如下:
import pandas as pd import numpy as np # 你的原始列表 lst = [ 'name, age, sex, height, weight', 'underweight,overweight,normal', 'David, 22, M, 185, -,-,78', 'Lily, 18, F, 165,-,75,-' ] # 1. 拼接最终表头 base_cols = [col.strip() for col in lst[0].split(',')] weight_tags = [tag.strip() for tag in lst[1].split(',')] # 替换掉原表头的weight字段,拼接三个体重分类子列 final_cols = base_cols[:-1] + weight_tags # 2. 清洗业务数据 processed_data = [] for row in lst[2:]: row_cols = [col.strip() for col in row.split(',')] # 将'-'替换为空值 cleaned_row = [np.nan if val == '-' else val for val in row_cols] processed_data.append(cleaned_row) # 3. 生成DataFrame,可选转换数值类型 df = pd.DataFrame(processed_data, columns=final_cols) df = df.apply(pd.to_numeric, errors='ignore')
运行后输出的df结构如下:
| name | age | sex | height | underweight | overweight | normal |
|---|---|---|---|---|---|---|
| David | 22 | M | 185 | NaN | NaN | 78 |
| Lily | 18 | F | 165 | NaN | 75 | NaN |
注意事项
- 如果你的业务数据行存在逗号数量不一致的情况,可以提前加长度校验逻辑过滤异常行
- 若需要保留'-'作为空值标识,删除替换
np.nan的步骤即可
内容的提问来源于stack exchange,提问作者Tranquil Oshan
相关产品推荐
相关产品推荐

