Spark2.0.0中如何通过SparkSession读取PNG等图片生成DataFrame?
Hey there! Great question—Spark 2.0.0 doesn’t have built-in support for reading image files directly into a DataFrame, but there are two solid ways to make this work for your use case. Let’s walk through them:
1. Custom Binary File Reading (No Third-Party Libraries)
Spark 2.0.0 doesn’t include the binaryFile data source (that arrived in later versions), but you can use Hadoop’s underlying input formats to read images as binary data, then convert to a DataFrame.
Step 1: Read images as binary RDD
First, create an RDD where each entry is a tuple of the file path and raw binary image content. We’ll use Hadoop’s SequenceFileInputFormat (adjust to FixedLengthInputFormat if you need fixed-length files):
from pyspark import SparkContext from org.apache.hadoop.io import BytesWritable, Text from org.apache.hadoop.mapreduce.lib.input import SequenceFileInputFormat # Configure Hadoop to read recursively if your images are in subfolders sc = spark.sparkContext sc._jsc.hadoopConfiguration().set("mapreduce.input.fileinputformat.input.dir.recursive", "true") # Read image files from local/HDFS path image_rdd = sc.newAPIHadoopFile( "/path/to/images", # Replace with your actual path SequenceFileInputFormat, Text, BytesWritable ).map(lambda x: (str(x[0]), bytearray(x[1].getBytes())))
Step 2: Convert RDD to DataFrame
Turn the binary RDD into a structured DataFrame with columns for file path and image data:
image_df = image_rdd.toDF(["file_path", "image_data"])
2. Using Databricks Spark-Image Library (Simpler Method)
Databricks maintains a spark-image library that adds image-specific features to Spark. For Spark 2.0.0, you can use a compatible version to read images directly into a DataFrame with useful metadata (like height, width, image format).
Step 1: Add the library
If you’re starting pyspark from the command line, include the dependency like this:
pyspark --packages com.databricks:spark-image_2.11:0.1.0
If you’re using a notebook, configure your cluster to include this Maven coordinate in its dependencies.
Step 2: Read images into DataFrame
Once the library is loaded, use the image data source to read your images:
image_df = spark.read.format("image") \ .option("pathGlobFilter", "*.png,*.jpg") # Filter to only include image files .load("/path/to/images")
This DataFrame will have an image struct column containing fields like origin (file path), height, width, nChannels, mode, and data (binary image content).
Saving DataFrame as HDFS SequenceFile
Once you have your image DataFrame, you can save it as a SequenceFile by converting it back to an RDD of Hadoop writable types, then using saveAsSequenceFile:
Option 1: Save from custom binary DataFrame
If you used the first approach, convert the DataFrame rows to Hadoop-compatible key-value pairs:
def row_to_writable(row): path = Text(row.file_path) data = BytesWritable(bytearray(row.image_data)) return (path, data) image_df.rdd.map(row_to_writable).saveAsSequenceFile("hdfs://your-hdfs-path/images-sequencefile")
Option 2: Save from spark-image DataFrame
If you used the spark-image library, extract the relevant fields from the struct column:
def image_struct_to_writable(row): image = row.image path = Text(image.origin) data = BytesWritable(bytearray(image.data)) return (path, data) image_df.rdd.map(image_struct_to_writable).saveAsSequenceFile("hdfs://your-hdfs-path/images-sequencefile")
Quick Notes
- If you can’t add third-party libraries, the custom Hadoop input method is your best bet—it’s reliable and doesn’t require extra dependencies.
- When working with HDFS, make sure your Spark cluster has the right permissions to read from the source path and write to the destination.
- SequenceFiles use a key-value structure, so we’re using the file path as the key and binary image data as the value, which is a standard pattern for this format.
内容的提问来源于stack exchange,提问作者skill_seeker

