基于JavaRDD<String>的Spark字符串排序及文件导出问题求助
Fixing Spark JavaRDD String Sorting & Export Issues
Hey there! Let's get your Spark string sorting and export working properly. Your code snippet cuts off early, but I can spot a few common missteps, plus I’ll share a complete, tested solution.
Common Issues in Your Partial Code
- Unnecessary Hadoop Configs: Those
fs.hdfs.imp(typo alert! Should befs.hdfs.impl) andfs.file.implsettings are almost never needed—Spark automatically handles local/HDFS file system detection based on your input/output paths. Manually setting these can cause unexpected errors. - Incomplete Logic: Your code stops before the critical sorting and export steps, which are the core of what you’re trying to do.
- Syntax Error: The
JavaRDD&l...snippet has a formatting typo—it should beJavaRDD<String>for proper generic typing.
Complete Working Solution
Here’s a full, runnable Java Spark program that reads a text file, sorts the strings, and exports the result:
import org.apache.spark.SparkConf; import org.apache.spark.api.java.JavaRDD; import org.apache.spark.api.java.JavaSparkContext; public class SparkStringSort { public static void main(String[] args) { // Initialize Spark configuration SparkConf conf = new SparkConf() .setMaster("local[*]") // Use all available cores for local testing .setAppName("Spark Sort"); // Use try-with-resources to auto-close SparkContext try (JavaSparkContext sparkContext = new JavaSparkContext(conf)) { // Read input text file into JavaRDD<String> JavaRDD<String> inputRDD = sparkContext.textFile("path/to/your/input/file.txt"); // Sort the strings: // - First arg: the key to sort by (the string itself) // - Second arg: true for ascending order, false for descending // - Third arg: number of partitions (adjust based on your data size) JavaRDD<String> sortedRDD = inputRDD.sortBy(s -> s, true, 1); // Export sorted RDD to a new directory (Spark will create this) // Note: The output directory MUST NOT exist beforehand, or Spark will throw an error sortedRDD.saveAsTextFile("path/to/your/output/sorted_files"); } } }
Key Notes for Success
- Spark Version Compatibility: If you’re using Spark 2.0 or later, it’s recommended to use
SparkSessioninstead ofJavaSparkContextfor a more modern API. Here’s a quick snippet for that:import org.apache.spark.sql.SparkSession; SparkSession spark = SparkSession.builder() .master("local[*]") .appName("Spark Sort") .getOrCreate(); JavaRDD<String> inputRDD = spark.read().textFile("input/path").javaRDD(); // Rest of the sorting/export logic stays the same - Output Directory Rule: Spark requires the output directory to not exist. If you run the program multiple times, delete the old output folder first or use a dynamic directory name.
- Custom Sorting: The default
sortByuses lexicographical order. For case-insensitive sorting, modify the lambda tos -> s.toLowerCase().
内容的提问来源于stack exchange,提问作者jlarrieux
相关产品推荐
相关产品推荐

