You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark DataFrame写入JSON与Parquet的行为差异咨询

Spark JSON vs Parquet Write Differences on AWS EMR 5.10 (Spark 2.2)

Hey there! Let's walk through the most common behavioral differences you might see when writing JSON vs Parquet to S3 with Spark 2.2 on EMR 5.10, plus how to adjust them if needed.

1. Default Compression Settings

This is one of the most noticeable immediate differences:

  • JSON writes: By default, Spark doesn’t apply compression to JSON output. Your part-000xx files are uncompressed plaintext, leading to larger file sizes and higher S3 storage usage.
    • To enable compression, set the relevant config before writing:
      spark.sql("SET spark.sql.json.compression.codec=snappy") // Options: snappy, gzip, deflate
      val data1 = sql("select * from csv_table")
      data1.write.json("s3://sparktest/jsonout/")
      
  • Parquet writes: EMR 5.10’s Spark 2.2 defaults to Snappy compression for Parquet (controlled by spark.sql.parquet.compression.codec=snappy). Combined with Parquet’s columnar format, this results in drastically smaller output files that are also faster to read later.

2. File Count & Size Behavior

  • JSON: Spark writes one part-000xx file per task in your job. If your CSV table is large and you have many tasks, you’ll end up with dozens or hundreds of small JSON files. To reduce file count, use coalesce() or repartition() to consolidate data before writing:
    val data1 = sql("select * from csv_table")
    data1.coalesce(3) // Merges data into 3 files (use repartition if full shuffling is needed)
         .write.json("s3://sparktest/jsonout/")
    
  • Parquet: Spark has built-in optimizations for Parquet to avoid excessive small files. It automatically adjusts file sizes based on data volume, and can merge small files during write (depending on cluster config). You’ll typically see fewer, larger Parquet files compared to unoptimized JSON writes.

3. Schema Handling

  • JSON: Schema merging is disabled by default in Spark 2.2. If your CSV table has evolving schemas or you’re appending data, you’ll need to explicitly enable it:
    spark.sql("SET spark.sql.json.mergeSchema=true")
    val data1 = sql("select * from csv_table")
    data1.write.mode("append").json("s3://sparktest/jsonout/")
    
  • Parquet: Schema merging is also disabled by default, but Parquet’s columnar format makes schema evolution more robust. Enable it with spark.sql.parquet.mergeSchema=true, and Spark will handle adding new columns to existing Parquet datasets more gracefully than JSON.

4. Write & Read Performance

  • JSON writes have lower serialization overhead (row-based plaintext) but slower IO due to larger uncompressed files.
  • Parquet has higher serialization overhead (converting rows to columns) but much faster IO thanks to compression and columnar storage. For large datasets, Parquet is almost always faster overall for writes, and significantly faster for subsequent reads.

内容的提问来源于stack exchange,提问作者kalyan chakravarthy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:22:29