You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将带表头、体重分类的字符串列表转换为规范pandas DataFrame

解决方案

核心思路是先单独处理表头拼接,再批量清洗业务数据后生成DataFrame,完整可运行代码如下:

import pandas as pd
import numpy as np

# 你的原始列表
lst = [
    'name, age, sex, height, weight',
    'underweight,overweight,normal',
    'David, 22, M, 185, -,-,78',
    'Lily, 18, F, 165,-,75,-'
]

# 1. 拼接最终表头
base_cols = [col.strip() for col in lst[0].split(',')]
weight_tags = [tag.strip() for tag in lst[1].split(',')]
# 替换掉原表头的weight字段,拼接三个体重分类子列
final_cols = base_cols[:-1] + weight_tags

# 2. 清洗业务数据
processed_data = []
for row in lst[2:]:
    row_cols = [col.strip() for col in row.split(',')]
    # 将'-'替换为空值
    cleaned_row = [np.nan if val == '-' else val for val in row_cols]
    processed_data.append(cleaned_row)

# 3. 生成DataFrame,可选转换数值类型
df = pd.DataFrame(processed_data, columns=final_cols)
df = df.apply(pd.to_numeric, errors='ignore')

运行后输出的df结构如下:

nameagesexheightunderweightoverweightnormal
David22M185NaNNaN78
Lily18F165NaN75NaN

注意事项

  • 如果你的业务数据行存在逗号数量不一致的情况,可以提前加长度校验逻辑过滤异常行
  • 若需要保留'-'作为空值标识,删除替换np.nan的步骤即可

内容的提问来源于stack exchange,提问作者Tranquil Oshan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 12:54:04