如何处理Pandas DataFrame空值并生成指定格式的数组新列
统一Pandas DataFrame中cve列的格式
原始数据与需求说明
给定如下Pandas DataFrame:
import pandas as pd data = [{'cve': [{'@href': 'https://ubuntu.com/security/CVE','@cvss_score':'10'}],'@href': '','@cvss_score': ''}, {'cve': '','@href': 'https://ubuntu.com/security/CVE','@cvss_score': '10'}, {'cve': '','@href': 'https://ubuntu.com/security/CVE','@cvss_score': '10'}, {'cve': [{'@href': 'https://ubuntu.com/security/CVE','@cvss_score':'10'}],'@href': '','@cvss_score': ''}] df = pd.DataFrame(data)
需求:当cve列为空字符串时,检查@href列是否有有效内容;若@href非空,则将@href和@cvss_score字段组合成与cve列格式一致的数组(如[{'@href': 'https://ubuntu.com/security/CVE','@cvss_score':'10'}]),并存入新列。
解决方案
通过自定义函数配合apply方法逐行处理数据,生成标准化后的新列:
# 定义格式标准化函数 def normalize_cve(row): if row['cve'] == '': if row['@href'] != '': return [{'@href': row['@href'], '@cvss_score': row['@cvss_score']}] return row['cve'] return row['cve'] # 生成新列 df['normalized_cve'] = df.apply(normalize_cve, axis=1) # 查看处理结果 print(df)
输出效果
处理后的DataFrame新增normalized_cve列,空cve行已按要求转换格式:
cve ... normalized_cve 0 [{'@href': 'https://ubuntu.com/security/CVE', ... ... [{'@href': 'https://ubuntu.com/security/CVE', ... 1 ... ... [{'@href': 'https://ubuntu.com/security/CVE', ... 2 ... ... [{'@href': 'https://ubuntu.com/security/CVE', ... 3 [{'@href': 'https://ubuntu.com/security/CVE', ... ... [{'@href': 'https://ubuntu.com/security/CVE', ... [4 rows x 4 columns]
内容的提问来源于stack exchange,提问作者Bandhala Raja Selvam
相关产品推荐
相关产品推荐

