You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark2.0.0中如何通过SparkSession读取PNG等图片生成DataFrame?

Reading Image Files (PNG/JPG) into DataFrame in Spark 2.0.0 & Saving as HDFS SequenceFile

Hey there! Great question—Spark 2.0.0 doesn’t have built-in support for reading image files directly into a DataFrame, but there are two solid ways to make this work for your use case. Let’s walk through them:

1. Custom Binary File Reading (No Third-Party Libraries)

Spark 2.0.0 doesn’t include the binaryFile data source (that arrived in later versions), but you can use Hadoop’s underlying input formats to read images as binary data, then convert to a DataFrame.

Step 1: Read images as binary RDD

First, create an RDD where each entry is a tuple of the file path and raw binary image content. We’ll use Hadoop’s SequenceFileInputFormat (adjust to FixedLengthInputFormat if you need fixed-length files):

from pyspark import SparkContext
from org.apache.hadoop.io import BytesWritable, Text
from org.apache.hadoop.mapreduce.lib.input import SequenceFileInputFormat

# Configure Hadoop to read recursively if your images are in subfolders
sc = spark.sparkContext
sc._jsc.hadoopConfiguration().set("mapreduce.input.fileinputformat.input.dir.recursive", "true")

# Read image files from local/HDFS path
image_rdd = sc.newAPIHadoopFile(
    "/path/to/images",  # Replace with your actual path
    SequenceFileInputFormat,
    Text,
    BytesWritable
).map(lambda x: (str(x[0]), bytearray(x[1].getBytes())))

Step 2: Convert RDD to DataFrame

Turn the binary RDD into a structured DataFrame with columns for file path and image data:

image_df = image_rdd.toDF(["file_path", "image_data"])

2. Using Databricks Spark-Image Library (Simpler Method)

Databricks maintains a spark-image library that adds image-specific features to Spark. For Spark 2.0.0, you can use a compatible version to read images directly into a DataFrame with useful metadata (like height, width, image format).

Step 1: Add the library

If you’re starting pyspark from the command line, include the dependency like this:

pyspark --packages com.databricks:spark-image_2.11:0.1.0

If you’re using a notebook, configure your cluster to include this Maven coordinate in its dependencies.

Step 2: Read images into DataFrame

Once the library is loaded, use the image data source to read your images:

image_df = spark.read.format("image") \
    .option("pathGlobFilter", "*.png,*.jpg")  # Filter to only include image files
    .load("/path/to/images")

This DataFrame will have an image struct column containing fields like origin (file path), height, width, nChannels, mode, and data (binary image content).

Saving DataFrame as HDFS SequenceFile

Once you have your image DataFrame, you can save it as a SequenceFile by converting it back to an RDD of Hadoop writable types, then using saveAsSequenceFile:

Option 1: Save from custom binary DataFrame

If you used the first approach, convert the DataFrame rows to Hadoop-compatible key-value pairs:

def row_to_writable(row):
    path = Text(row.file_path)
    data = BytesWritable(bytearray(row.image_data))
    return (path, data)

image_df.rdd.map(row_to_writable).saveAsSequenceFile("hdfs://your-hdfs-path/images-sequencefile")

Option 2: Save from spark-image DataFrame

If you used the spark-image library, extract the relevant fields from the struct column:

def image_struct_to_writable(row):
    image = row.image
    path = Text(image.origin)
    data = BytesWritable(bytearray(image.data))
    return (path, data)

image_df.rdd.map(image_struct_to_writable).saveAsSequenceFile("hdfs://your-hdfs-path/images-sequencefile")

Quick Notes

  • If you can’t add third-party libraries, the custom Hadoop input method is your best bet—it’s reliable and doesn’t require extra dependencies.
  • When working with HDFS, make sure your Spark cluster has the right permissions to read from the source path and write to the destination.
  • SequenceFiles use a key-value structure, so we’re using the file path as the key and binary image data as the value, which is a standard pattern for this format.

内容的提问来源于stack exchange,提问作者skill_seeker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:42:16