You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python DataFrame列筛选:按指定列列表保留列并转JSON行

解决方案

步骤说明

  1. 先筛选出需要保留的列:原DataFrame的Name列(匹配预期输出需求),加上列列表col_list中存在的Country列。
  2. 逐行处理数据,自动移除值为NaN的字段。
  3. 针对特定分组标题行(如包含"through"的行),将键名从Name替换为col_list中的Name of Company,对齐预期输出格式。

代码实现

import pandas as pd
import numpy as np
import json

# 构造示例DataFrame
data = [
    ["Directly held interests", np.nan, np.nan, np.nan, np.nan],
    ["Ferrari S.p.A.", "Italy", "Manufacturing", "100%", "-%"],
    ["Indirectly held through Ferrari S.p.A.", np.nan, np.nan, np.nan, np.nan]
]
columns = ["Name", "Country", "Nature of business", "Shares held by the Group", "Shares held by NCI"]
df2 = pd.DataFrame(data, columns=columns)

# 给定的目标列列表
col_list = ['Name of Company','Incorporated','Country']

# 筛选要保留的列:Name列 + 列表中存在的列
existing_cols = [col for col in col_list if col in df2.columns]
selected_cols = ['Name'] + existing_cols
df_selected = df2[selected_cols]

# 逐行生成符合要求的JSON
for _, row in df_selected.iterrows():
    # 移除NaN值并转为字典
    clean_row = row.dropna().to_dict()
    # 对间接持有分组的标题行替换键名
    if len(clean_row) == 1 and 'Name' in clean_row and 'through' in clean_row['Name']:
        clean_row['Name of Company'] = clean_row.pop('Name')
    # 输出JSON字符串
    print(json.dumps(clean_row))

输出结果

运行代码后将得到预期输出:

{"Name":"Directly held interests"}
{"Name":"Ferrari S.p.A.","Country":"Italy"}
{"Name of Company":"Indirectly held through Ferrari S.p.A."}

灵活调整说明

如果实际数据中分组标题行的判断逻辑不同(比如不是通过"through"关键词),可以修改代码中的判断条件,例如根据行内非NaN字段数量、特定前缀等规则调整键名替换逻辑。

内容的提问来源于stack exchange,提问作者emiley mille

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 22:08:10