Spark 1.6读取带表头尾的CSV仅识别首列,如何配置参数?
Hey there, let's troubleshoot why your Spark job is only pulling in the first column of your CSV file. The root cause here is that you're not telling Spark how your data is separated, plus we need to handle the non-standard header and trailing row in your file.
What's Going Wrong?
By default, the Databricks CSV library (used in Spark 1.6) expects comma-separated values, but your data is clearly using spaces as delimiters. That's why every line is being treated as a single column value. Additionally, your file has a dummy first row (123456) and a trailing summary row (Count-21) that we need to exclude to get clean data.
Step-by-Step Solution
Here's the adjusted code with the necessary configurations to load all columns correctly:
// Fix the library name (you had a typo: databrics → databricks) val df = spark.read .format("com.databricks.spark.csv") .option("header", true) // Treat the second row as the actual header (Name, age, branch) .option("delimiter", " ") // Specify space as the column separator .option("skipRows", 1) // Skip the first dummy row ("123456") .option("inferSchema", false) // Keep schema as string (adjust if needed) .load(path) // Filter out the trailing "Count-21" row val cleanDf = df.filter(!$"Name".equalTo("Count-21")) // Now show all columns cleanDf.show()
Key Configurations Explained
delimiter: Tells Spark to split columns on spaces instead of the default commas. This is the critical fix for getting all columns to load.skipRows: Skips the first row (123456) since it's not a valid header for your data columns.header=true: Instructs Spark to use the second row (Name age branch) as the column names.- Filtering the trailing row: Spark 1.6 doesn't have a built-in parameter to skip the last row, so we filter out the row where the
Namecolumn equalsCount-21.
Quick Note
Double-check the library name: it's com.databricks.spark.csv (not com.databrics.csv — you had a typo there, which would cause a class not found error if left unfixed).
内容的提问来源于stack exchange,提问作者jaggu

