如何从pandas profiling中提取缺失值、唯一值等指标导出为Excel表格
实现方案
你需要的所有统计指标都存储在profset["variables"]下每个变量对应的属性字典中,直接遍历提取即可,完整可运行代码如下:
import pandas as pd # 旧版本pandas-profiling用下面的导入 from pandas_profiling import ProfileReport # 新版本ydata-profiling替换为下面的导入即可 # from ydata_profiling import ProfileReport # 你已有的原有逻辑 df1 = pd.read_excel('df.xlsx') profile = df1.profile_report(title="Dataset Profiling Report") profile.to_file('dataset_report.html') profset = profile.description_set attributes = profset["variables"] # 新增提取指标逻辑 stats_records = [] for attr_name, attr_detail in attributes.items(): stats_records.append({ "Attributes": attr_name, "n_Missing": attr_detail["n_missing"], "p_missing": attr_detail["p_missing"], "n_distinct": attr_detail["n_distinct"], "p_distinct": attr_detail["p_distinct"] }) # 转换为DataFrame并导出Excel stats_df = pd.DataFrame(stats_records) stats_df.to_excel("变量统计结果.xlsx", index=False)
- 若运行时提示键不存在,可先打印单变量的所有键确认字段名:
print(attributes[list(attributes.keys())[0]].keys()),不同版本的库字段名可能存在微小差异,对应替换即可。
内容的提问来源于stack exchange,提问作者GiveGet_15
相关产品推荐
相关产品推荐

