如何遍历CSV表头生成带空值检测的Python字典列表?
实现方案
以下是满足需求的Python代码,基于pandas完成字典列表的生成:
import pandas as pd # 替换为实际加载CSV的代码,比如 df = pd.read_csv('your_file.csv') data = { 'id': [1, 2, 3, 4], 'NAME': ['Henry', 'Joe', pd.NA, 'Smith'], 'location': ['London', 'Peru', 'Germany', pd.NA] } df = pd.DataFrame(data) result_list = [] for idx, col_name in enumerate(df.columns): # 构建基础字典 col_dict = { "item": col_name, "seq": (idx + 1) * 2, # 对应示例中seq的计算规则 "is_null": df[col_name].isnull().any(), "version": "done" } # 判断列名是否包含大写字母,添加new字段 if any(char.isupper() for char in col_name): col_dict["new"] = col_name.lower() result_list.append(col_dict) # 输出结果 print(result_list)
关键逻辑说明
- 遍历列名与索引:使用
enumerate(df.columns)同时获取列的索引和名称,用于计算seq值(示例中seq为索引+1后乘2,对应第一个列seq=2,第二个=4,以此类推)。 - 空值判断:
df[col_name].isnull().any()会快速检查列中是否存在任意空值(NaN/pd.NA),返回布尔值填充is_null字段。 - 大写字母检测:通过
any(char.isupper() for char in col_name)判断列名是否包含大写字母,满足条件则添加new字段存储小写形式。 - 固定字段:所有字典都强制加入
version: "done"字段。
运行上述代码后,输出结果与你提供的示例一致:
[{"item": "id", "is_null": false, "seq": 2, "version": "done"}, {"item": "NAME", "new": "name", "is_null": true, "seq": 4, "version": "done"}, {"item": "location", "is_null": true, "seq": 6, "version": "done"}]
内容的提问来源于stack exchange,提问作者Pikun95
相关产品推荐
相关产品推荐

