如何将Python中的ProfileReport导出为PDF或Excel格式?
导出Pandas Profiling的ProfileReport为PDF/Excel
针对200列的表格,以下是将Python生成的profile-report HTML结果导出为PDF(用于管理层讨论)和Excel的实用方案:
一、导出为PDF
方法1:使用WeasyPrint直接转换HTML(推荐)
WeasyPrint可直接解析HTML/CSS生成高质量PDF,适配静态内容转换,适合多列表格场景。
- 安装依赖:
pip install weasyprint pandas-profiling
- 代码实现:
from pandas_profiling import ProfileReport import pandas as pd from weasyprint import HTML # 加载你的200列表格数据 df = pd.read_csv("your_data.csv") # 生成ProfileReport并保存为HTML profile = ProfileReport(df, title="Data Profile Report") profile.to_file("profile_report.html") # 将HTML转换为PDF,设置横向页面适配多列 HTML("profile_report.html").write_pdf( "profile_report.pdf", stylesheets=[""" @page { size: landscape; margin: 1cm; } table { width: 100%; border-collapse: collapse; } th, td { border: 1px solid #ddd; padding: 4px; font-size: 8px; } """] )
注:自定义CSS可调整表格字体大小、页面边距,避免200列内容被截断。
方法2:使用Selenium渲染后导出PDF(适配动态内容)
如果HTML包含折叠面板等动态交互元素,Selenium可模拟浏览器完全渲染后再导出PDF。
- 安装依赖:
pip install selenium pandas-profiling
需额外下载对应浏览器驱动(如ChromeDriver)并配置到环境变量。
- 代码实现:
from pandas_profiling import ProfileReport import pandas as pd from selenium import webdriver from selenium.webdriver.chrome.options import Options df = pd.read_csv("your_data.csv") profile = ProfileReport(df, title="Data Profile Report") profile.to_file("profile_report.html") # 配置Chrome浏览器打印参数 chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") chrome_options.add_experimental_option("prefs", { "printing.print_preview_sticky_settings.appState": """{ "recentDestinations": [{"id": "Save as PDF", "origin": "local", "account": ""}], "selectedDestinationId": "Save as PDF", "version": 2, "isHeaderFooterEnabled": false, "pageSize": {"width": 11.69, "height": 8.27} # 横向A4 }""" }) chrome_options.add_argument("--kiosk-printing") # 启动浏览器并导出PDF driver = webdriver.Chrome(options=chrome_options) driver.get("file:///path/to/your/profile_report.html") driver.execute_script("window.print();") driver.quit()
二、导出为Excel
Excel适合提取可编辑的统计数据,可将ProfileReport核心内容拆分到不同工作表:
- 安装依赖:
pip install pandas-profiling openpyxl
- 代码实现:
from pandas_profiling import ProfileReport import pandas as pd df = pd.read_csv("your_data.csv") profile = ProfileReport(df, title="Data Profile Report") # 提取ProfileReport核心统计数据 description = profile.description_set # 创建Excel写入器,拆分内容到不同sheet with pd.ExcelWriter("profile_report.xlsx", engine="openpyxl") as writer: # 概述信息 pd.DataFrame([description["overview"]]).to_excel(writer, sheet_name="概述", index=False) # 200列变量详细统计 description["variables"].to_excel(writer, sheet_name="变量统计", index=True) # 相关性矩阵(若存在) if "correlations" in description: description["correlations"]["pearson"].to_excel(writer, sheet_name="相关性矩阵") # 缺失值统计 description["missing"].to_excel(writer, sheet_name="缺失值分析")
内容的提问来源于stack exchange,提问作者Natali
相关产品推荐
相关产品推荐

