Spark NLP运行PySpark报AnalysisException结构体字段0不存在错误

问题复现代码
运行如下PySpark代码统计标签分布时抛出异常:
conll_data.select(F.explode(F.arrays_zip('token.result','label.result')).alias("cols")) \ .select(F.expr("cols['0']").alias("token"), F.expr("cols['1']").alias("ground_truth"))\ .groupBy('ground_truth')\ .count()\ .orderBy('count', ascending=False)\ .show(100,truncate=False)
异常信息:
AnalysisException: No such struct field 0 in result, result
当前运行环境依赖:
jupyterlab SQLAlchemy==0.7.1 spark-nlp==3.4.4 pyspark==3.1.2 numpy== 1.19.2 pandas==1.3.2 openpyxl==3.0.9 jupyter_contrib_nbextensions spark-nlp-display pyarrow==3.0.0 streamlit==1.1.0 scipy==1.7.3 Tensorflow==2.5.0 tensorflow-addons python==3.7.4
报错原因
F.arrays_zip合并多个数组生成Struct数组时,会保留传入列的原名作为Struct内部的字段名,不会自动生成0、1这类数字索引字段。
传入的两个参数token.result、label.result提取后列名均为result,因此合并后cols结构体内只有两个重名的result字段,不存在名为0、1的字段,最终触发字段不存在的分析异常。
修复方法
最小改动方案:传入arrays_zip时先给两个列设置独立别名,后续直接按别名取字段即可,不需要用数字索引访问:
conll_data.select(F.explode(F.arrays_zip( F.col('token.result').alias('token'), F.col('label.result').alias('ground_truth') )).alias("cols")) \ .select("cols.token", "cols.ground_truth")\ .groupBy('ground_truth')\ .count()\ .orderBy('count', ascending=False)\ .show(100,truncate=False)
内容的提问来源于stack exchange,提问作者Tokvl
相关产品推荐
相关产品推荐

