Scala中如何使用.drop()方法删除CSV文件的首行表头?
问题解答
drop(1)作用说明
你修改后的代码中getLines.drop(1)会直接跳过整个CSV文件的首行,不会仅删除首行的第一个元素。Source.fromFile(file).getLines返回的是字符串迭代器,迭代器的每个元素对应文件中完整的一行内容,drop(n)方法的作用是跳过迭代器的前n个元素,对应就是跳过文件的前n行,所以drop(1)正好可以实现跳过首行表头的效果。
存为Seq/数组的实现方案
基础实现(仅读取行内容)
import scala.io.Source // 存为Seq[String],每个元素对应一行数据 val csvDataSeq: Seq[String] = Source.fromFile("你的CSV文件路径") .getLines .drop(1) // 跳过表头行 .toSeq // 存为数组的版本 val csvDataArray: Array[String] = Source.fromFile("你的CSV文件路径") .getLines .drop(1) .toArray
进阶实现(拆分字段转为结构化对象)
如果需要直接把每行内容拆分为对应字段,可参考以下写法:
import scala.io.Source // 定义和表头对应的结构化数据类 case class CsvRow( birthday: String, zipCode: String, name: String, city: String, state: String, country: String ) val structuredData: Seq[CsvRow] = Source.fromFile("你的CSV文件路径") .getLines .drop(1) .map(line => { val fields = line.split(",") CsvRow( birthday = fields(0), zipCode = fields(1), name = fields(2), city = fields(3), state = fields(4), country = fields(5) ) }) .toSeq
注意事项
- 生产环境使用建议补充资源关闭逻辑,避免文件句柄泄漏,可以在finally块中调用
Source.close()方法。 - 如果你的CSV包含被引号包裹、内部带逗号的特殊字段,不要直接用
split(",")拆分,建议使用专业CSV解析库处理,避免字段拆分错误。
内容的提问来源于stack exchange,提问作者Art smart
相关产品推荐
相关产品推荐

