在Apache Zeppelin中调用pandas_profiling.to_widgets()遇NotImplementedError
问题:Apache Zeppelin中使用pandas_profiling分析Spark DataFrame报错NotImplementedError
代码与环境信息
执行的PySpark代码:
%pyspark import sys print(sys.version_info) import numpy as np print("numpy: ", np.__version__) import pandas as pd print("pandas: ", pd.__version__) import pandas_profiling as pp print("pandas_profiling: ", pp.__version__) from pandas_profiling import ProfileReport df = spark.sql("SELECT * FROM database.table") profile = ProfileReport(df, title="Report: table") profile.to_widgets()
环境版本输出:
sys.version_info(major=3, minor=6, micro=8, releaselevel='final', serial=0) numpy: 1.19.5 pandas: 1.1.5 pandas_profiling: 3.1.0
报错栈信息:
Fail to execute line 19: profile.to_widgets() Traceback (most recent call last): File "/tmp/1662648724242-0/zeppelin_python.py", line 158, in <module> exec(code, _zcUserQueryNameSpace) File "<stdin>", line 19, in <module> File "/usr/local/lib/python3.6/site-packages/pandas_profiling/profile_report.py", line 414, in to_widgets display(self.widgets) File "/usr/local/lib/python3.6/site-packages/pandas_profiling/profile_report.py", line 197, in widgets self._widgets = self._render_widgets() File "/usr/local/lib/python3.6/site-packages/pandas_profiling/profile_report.py", line 315, in _render_widgets report = self.report File "/usr/local/lib/python3.6/site-packages/pandas_profiling/profile_report.py", line 179, in report self._report = get_report_structure(self.config, self.description_set) File "/usr/local/lib/python3.6/site-packages/pandas_profiling/profile_report.py", line 166, in description_set self._sample, File "/usr/local/lib/python3.6/site-packages/pandas_profiling/model/describe.py", line 56, in describe check_dataframe(df) File "/usr/local/lib/python3.6/site-packages/multimethod/__init__.py", line 209, in __call__ return self[tuple(map(self.get_type, args))](*args, **kwargs) File "/usr/local/lib/python3.6/site-packages/pandas_profiling/model/dataframe.py", line 10, in check_dataframe raise NotImplementedError() NotImplementedError
解决方法
核心原因
pandas_profiling的ProfileReport仅支持pandas DataFrame,不直接兼容Spark DataFrame,传入Spark对象会触发数据类型检查失败,抛出NotImplementedError。同时Apache Zeppelin对Jupyter风格的widgets支持有限,to_widgets()方法本身也无法在Zeppelin环境中正常工作。
可行解决方案(Zeppelin内部解决)
- 转换Spark DataFrame为pandas DataFrame
- 若数据量较小,直接用
toPandas()转换;若数据量较大,先采样避免内存溢出。
- 若数据量较小,直接用
- 生成HTML报告并在Zeppelin中展示
- 替换
profile.to_widgets()为生成HTML报告,再通过Zeppelin的display.html()方法展示。
- 替换
修改后的代码示例:
%pyspark import sys print(sys.version_info) import numpy as np print("numpy: ", np.__version__) import pandas as pd print("pandas: ", pd.__version__) import pandas_profiling as pp print("pandas_profiling: ", pp.__version__) from pandas_profiling import ProfileReport # 读取Spark表数据 spark_df = spark.sql("SELECT * FROM database.table") # 转换为pandas DataFrame(大数据量建议先采样) # 采样示例:spark_df = spark_df.sample(fraction=0.1, seed=42) pandas_df = spark_df.toPandas() # 生成分析报告 profile = ProfileReport(pandas_df, title="Report: table") # 生成HTML内容并在Zeppelin中展示 html_report = profile.to_html() display.html(html_report, width="100%", height="800px")
注意事项
- 若原表数据量极大,直接转换会导致Executor内存溢出,必须先通过
sample()或limit()减少数据量。 - 确保Zeppelin的Python环境有足够内存处理转换后的pandas DataFrame。
内容的提问来源于stack exchange,提问作者Lucas Soares
相关产品推荐
相关产品推荐

