You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark NLP运行PySpark报AnalysisException结构体字段0不存在错误

问题截图

问题复现代码

运行如下PySpark代码统计标签分布时抛出异常:

conll_data.select(F.explode(F.arrays_zip('token.result','label.result')).alias("cols")) \
          .select(F.expr("cols['0']").alias("token"),
                  F.expr("cols['1']").alias("ground_truth"))\
          .groupBy('ground_truth')\
          .count()\
          .orderBy('count', ascending=False)\
          .show(100,truncate=False)

异常信息:

AnalysisException: No such struct field 0 in result, result

当前运行环境依赖:

jupyterlab
SQLAlchemy==0.7.1
spark-nlp==3.4.4
pyspark==3.1.2
numpy== 1.19.2
pandas==1.3.2
openpyxl==3.0.9
jupyter_contrib_nbextensions
spark-nlp-display
pyarrow==3.0.0
streamlit==1.1.0
scipy==1.7.3
Tensorflow==2.5.0
tensorflow-addons
python==3.7.4

报错原因

F.arrays_zip合并多个数组生成Struct数组时,会保留传入列的原名作为Struct内部的字段名,不会自动生成0、1这类数字索引字段。
传入的两个参数token.result、label.result提取后列名均为result,因此合并后cols结构体内只有两个重名的result字段,不存在名为0、1的字段,最终触发字段不存在的分析异常。


修复方法

最小改动方案:传入arrays_zip时先给两个列设置独立别名,后续直接按别名取字段即可,不需要用数字索引访问:

conll_data.select(F.explode(F.arrays_zip(
    F.col('token.result').alias('token'),
    F.col('label.result').alias('ground_truth')
)).alias("cols")) \
          .select("cols.token", "cols.ground_truth")\
          .groupBy('ground_truth')\
          .count()\
          .orderBy('count', ascending=False)\
          .show(100,truncate=False)

内容的提问来源于stack exchange,提问作者Tokvl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.02 09:54:29