Spark DataFrame写入JSON与Parquet的行为差异咨询
Spark JSON vs Parquet Write Differences on AWS EMR 5.10 (Spark 2.2)
Hey there! Let's walk through the most common behavioral differences you might see when writing JSON vs Parquet to S3 with Spark 2.2 on EMR 5.10, plus how to adjust them if needed.
1. Default Compression Settings
This is one of the most noticeable immediate differences:
- JSON writes: By default, Spark doesn’t apply compression to JSON output. Your
part-000xxfiles are uncompressed plaintext, leading to larger file sizes and higher S3 storage usage.- To enable compression, set the relevant config before writing:
spark.sql("SET spark.sql.json.compression.codec=snappy") // Options: snappy, gzip, deflate val data1 = sql("select * from csv_table") data1.write.json("s3://sparktest/jsonout/")
- To enable compression, set the relevant config before writing:
- Parquet writes: EMR 5.10’s Spark 2.2 defaults to Snappy compression for Parquet (controlled by
spark.sql.parquet.compression.codec=snappy). Combined with Parquet’s columnar format, this results in drastically smaller output files that are also faster to read later.
2. File Count & Size Behavior
- JSON: Spark writes one
part-000xxfile per task in your job. If your CSV table is large and you have many tasks, you’ll end up with dozens or hundreds of small JSON files. To reduce file count, usecoalesce()orrepartition()to consolidate data before writing:val data1 = sql("select * from csv_table") data1.coalesce(3) // Merges data into 3 files (use repartition if full shuffling is needed) .write.json("s3://sparktest/jsonout/") - Parquet: Spark has built-in optimizations for Parquet to avoid excessive small files. It automatically adjusts file sizes based on data volume, and can merge small files during write (depending on cluster config). You’ll typically see fewer, larger Parquet files compared to unoptimized JSON writes.
3. Schema Handling
- JSON: Schema merging is disabled by default in Spark 2.2. If your CSV table has evolving schemas or you’re appending data, you’ll need to explicitly enable it:
spark.sql("SET spark.sql.json.mergeSchema=true") val data1 = sql("select * from csv_table") data1.write.mode("append").json("s3://sparktest/jsonout/") - Parquet: Schema merging is also disabled by default, but Parquet’s columnar format makes schema evolution more robust. Enable it with
spark.sql.parquet.mergeSchema=true, and Spark will handle adding new columns to existing Parquet datasets more gracefully than JSON.
4. Write & Read Performance
- JSON writes have lower serialization overhead (row-based plaintext) but slower IO due to larger uncompressed files.
- Parquet has higher serialization overhead (converting rows to columns) but much faster IO thanks to compression and columnar storage. For large datasets, Parquet is almost always faster overall for writes, and significantly faster for subsequent reads.
内容的提问来源于stack exchange,提问作者kalyan chakravarthy
相关产品推荐
相关产品推荐

