You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将PySpark DataFrame单行中的bytearray转换为单字节多行列

实现方案

你之前直接用explode失败的原因是Spark的explode仅支持数组、映射类型,Binary类型不属于可直接拆分的集合类型,需要先转换为数组结构再拆分。

核心逻辑是先通过UDF将Binary类型的字节序列转换为存储各字节十进制值的数组,再调用explode拆分为多行即可,具体实现代码如下:

from pyspark.sql import functions as F
from pyspark.sql.types import ArrayType, IntegerType

# 定义UDF:将二进制内容转换为各字节的十进制值数组
binary_to_int_array = F.udf(lambda byte_data: [x for x in byte_data], ArrayType(IntegerType()))

# 示例原始数据,可替换为你通过spark.read.format('binaryfile').load得到的DataFrame
import pandas as pd
test_df = pd.DataFrame({'content': [bytearray(b'\x01%\xeb\x8cH\x89')]})
spark_df = spark.createDataFrame(test_df)

# 拆分字节为单独行
result_df = spark_df.select(
    F.explode(binary_to_int_array(F.col("content"))).alias("content")
)

# 输出结果
result_df.show()

运行后得到的结果和预期完全一致:

+-------+
|content|
+-------+
|      1|
|     37|
|    235|
|    140|
|     72|
|    137|
+-------+

内容的提问来源于stack exchange,提问作者pmdaly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 22:45:06