Spark RDD词频统计示例编译失败,请求技术支持
Spark RDD单词计数示例编译错误解决
问题描述
我是Scala和Spark新手,尝试运行Spark官方的RDD单词计数示例(统计文本文件中每个单词的出现次数)但未成功。将教授提供的可运行Scala模板代码与官方示例整合后,出现编译错误。
错误代码
// Program assumes that // 1) NO folder named "output" exists in the same directory before execution, and // 2) the file "some_words.txt" already exists in the same directory before execution // Calculates the frequency of words that occur in a document, outputting duples like "(Anthill, 3)" import org.apache.spark.SparkContext import org.apache.spark.SparkContext._ import org.apache.spark.SparkConf object SimpleApp { def main(args: Array[String]) { val fileToRead = "input.txt" val conf = new SparkConf().setAppName("Simple Application") val sc = new SparkContext(conf) text_file = spark.sparkContext.textFile("some_words.txt") counts = ( text_file.flatMap(lambda line: line.split(" ")) .map(lambda word: (word, 1)) .reduceByKey(lambda a, b: a + b) ) println(counts) // Folder to save to wordsWithCount.saveAsTextFile("output") sc.stop() } }
报错信息
')' expected but '(' found. [error] fileData.flatMap(lambda line: line.split(" ")) ^ identifier expected but integer literal found. [error] .map(lambda word: (word,1)) ^ ')' expected but '}' found. [error] }
问题分析
编译错误的核心原因是混用Python语法与Scala语法,同时存在变量声明、命名不一致问题:
- Scala无
lambda关键字,匿名函数需用=>语法实现 - Scala是强类型语言,变量必须用
val/var声明,不能直接赋值 - 变量名混乱:
text_file未声明、wordsWithCount未定义,与前面的counts不匹配 - 错误调用
spark.sparkContext,当前代码使用SparkContext实例sc,应直接调用sc.textFile
修正后的代码
// Program assumes that // 1) NO folder named "output" exists in the same directory before execution, and // 2) the file "some_words.txt" already exists in the same directory before execution // Calculates the frequency of words that occur in a document, outputting tuples like "(Anthill, 3)" import org.apache.spark.SparkContext import org.apache.spark.SparkContext._ import org.apache.spark.SparkConf object SimpleApp { def main(args: Array[String]) { val conf = new SparkConf().setAppName("Word Count Application") val sc = new SparkContext(conf) // 读取文本文件 val textFile = sc.textFile("some_words.txt") // 执行单词计数逻辑 val wordCounts = textFile .flatMap(line => line.split(" ")) // Scala匿名函数替代Python lambda .map(word => (word, 1)) .reduceByKey((a, b) => a + b) // 打印计数结果 wordCounts.foreach(println) // 保存结果到output目录 wordCounts.saveAsTextFile("output") sc.stop() } }
构建脚本
sbt package spark-submit --master local[4] target/scala-2.11/simple-project_2.11-1.0.jar
内容的提问来源于stack exchange,提问作者Stev
相关产品推荐
相关产品推荐

