Python DataFrame列筛选:按指定列列表保留列并转JSON行
解决方案
步骤说明
- 先筛选出需要保留的列:原DataFrame的
Name列(匹配预期输出需求),加上列列表col_list中存在的Country列。 - 逐行处理数据,自动移除值为
NaN的字段。 - 针对特定分组标题行(如包含"through"的行),将键名从
Name替换为col_list中的Name of Company,对齐预期输出格式。
代码实现
import pandas as pd import numpy as np import json # 构造示例DataFrame data = [ ["Directly held interests", np.nan, np.nan, np.nan, np.nan], ["Ferrari S.p.A.", "Italy", "Manufacturing", "100%", "-%"], ["Indirectly held through Ferrari S.p.A.", np.nan, np.nan, np.nan, np.nan] ] columns = ["Name", "Country", "Nature of business", "Shares held by the Group", "Shares held by NCI"] df2 = pd.DataFrame(data, columns=columns) # 给定的目标列列表 col_list = ['Name of Company','Incorporated','Country'] # 筛选要保留的列:Name列 + 列表中存在的列 existing_cols = [col for col in col_list if col in df2.columns] selected_cols = ['Name'] + existing_cols df_selected = df2[selected_cols] # 逐行生成符合要求的JSON for _, row in df_selected.iterrows(): # 移除NaN值并转为字典 clean_row = row.dropna().to_dict() # 对间接持有分组的标题行替换键名 if len(clean_row) == 1 and 'Name' in clean_row and 'through' in clean_row['Name']: clean_row['Name of Company'] = clean_row.pop('Name') # 输出JSON字符串 print(json.dumps(clean_row))
输出结果
运行代码后将得到预期输出:
{"Name":"Directly held interests"} {"Name":"Ferrari S.p.A.","Country":"Italy"} {"Name of Company":"Indirectly held through Ferrari S.p.A."}
灵活调整说明
如果实际数据中分组标题行的判断逻辑不同(比如不是通过"through"关键词),可以修改代码中的判断条件,例如根据行内非NaN字段数量、特定前缀等规则调整键名替换逻辑。
内容的提问来源于stack exchange,提问作者emiley mille
相关产品推荐
相关产品推荐

