Spark DataFrame写入可读JSON遇序列化兼容问题,求替代方案
That InvalidClassException you're hitting comes down to a version mismatch of the org.apache.commons.lang3.time.FastDateFormat class between your local build and the cluster. The serialVersionUID in the version you packaged doesn't match what's running on the cluster, which breaks serialization when Spark tries to distribute your job.
Here are practical alternatives to get readable ASCII JSON output using Spark's APIs, without running into that error:
1. Use Spark's Native write.json with Readable Configuration
Skip toJSON entirely and leverage Spark's built-in JSON writer, which lets you enable human-readable output and specify ASCII encoding directly:
df.write .option("multiLine", "true") // Outputs pretty-printed, readable JSON .option("encoding", "US-ASCII") // Forces ASCII encoding for all output .json("/your/output/path")
This approach avoids the dependency on FastDateFormat that toJSON introduces, so serialization mismatches won't be an issue. The output will be clean, formatted JSON files that are easy to inspect.
2. Resolve the Dependency Conflict (If You Need toJSON)
If you have to stick with dataframe.toJSON.write.save, fix the commons-lang3 version mismatch:
- For SBT projects: Mark commons-lang3 as
Providedso your build doesn't package it, and uses the cluster's installed version instead:libraryDependencies += "org.apache.commons" % "commons-lang3" % "3.12.0" % Provided - For Maven projects: Exclude commons-lang3 from your Spark dependencies to avoid packaging a conflicting version:
<dependency> <groupId>org.apache.spark</groupId> <artifactId>spark-sql_2.12</artifactId> <version>3.3.0</version> <scope>provided</scope> <exclusions> <exclusion> <groupId>org.apache.commons</groupId> <artifactId>commons-lang3</artifactId> </exclusion> </exclusions> </dependency>
Just make sure the version you reference (if any) matches the one running on your cluster.
3. Use Spark's to_json Function for Manual JSON Conversion
Another reliable method is to use Spark's native to_json function to convert rows to JSON strings, then save as text files with ASCII encoding:
import org.apache.spark.sql.functions.{to_json, struct} // Convert all columns into a single JSON string column val jsonStringDF = df.select(to_json(struct(df.columns.map(col): _*)).alias("json_content")) // Save as text files with enforced ASCII encoding jsonStringDF.write .option("encoding", "US-ASCII") .text("/your/output/path")
This uses Spark's internal JSON handling, so no external dependency issues. The output will be one JSON object per line—both readable and easy to process downstream.
内容的提问来源于stack exchange,提问作者user1124702

