使用Circe解码JSON行文件,如何过滤country为us的Person实例?
处理大JSON行文件并过滤Person实例的最佳实现
首先先优化你现有代码里的几个问题:
Decoder[Person]不需要在map循环里重复创建,放在全局作用域或方法内部即可,避免重复初始化- 超大文件不要用
toList把所有行一次性加载到内存,直接用getLines返回的Iterator流式处理,内存更友好
下面是几种符合你需求的实现方式:
方式一:Option结合flatMap(推荐)
这种方式逻辑清晰,流式处理,只保留符合条件的Person实例,同时可灵活处理解码失败的行:
import io.circe.Decoder import io.circe.generic.semiauto.deriveDecoder import io.circe.parser.decode import scala.io.Source case class Person(name: String, age: Int, country: String) // 全局定义Decoder,避免重复创建 implicit val personDecoder: Decoder[Person] = deriveDecoder[Person] val personList: List[Person] = Source.fromResource("Persons.json") .getLines .flatMap { line => decode[Person](line) match { // 解码成功且country为"us",返回Some实例 case Right(p @ Person(_, _, "us")) => Some(p) // 解码成功但country不是"us",返回None被flatMap自动过滤 case Right(_) => None // 解码失败,这里选择抛出异常,也可改成None忽略错误行 case Left(ex) => throw new RuntimeException(ex) } } .toList // 最后转成List,若不需要存储全部数据,可保持Iterator流式处理
方式二:嵌套模式匹配
和方式一逻辑类似,仅用嵌套匹配实现过滤,但会对同一行调用两次decode,效率略低:
import io.circe.Decoder import io.circe.generic.semiauto.deriveDecoder import io.circe.parser.decode import scala.io.Source case class Person(name: String, age: Int, country: String) implicit val personDecoder: Decoder[Person] = deriveDecoder[Person] val personList: List[Person] = Source.fromResource("Persons.json") .getLines .collect { // collect只保留匹配成功的元素 case line => decode[Person](line) match { case Right(p @ Person(_, _, "us")) => p } // 可添加case分支处理解码失败的情况,比如抛出异常或忽略 } .toList
方式三:自定义Decoder在解码阶段过滤
把过滤逻辑整合到解码过程中,不符合条件的直接判定为解码失败:
import io.circe.Decoder import io.circe.generic.semiauto.deriveDecoder import io.circe.parser.decode import scala.io.Source case class Person(name: String, age: Int, country: String) // 自定义Decoder,仅country为"us"的实例能解码成功 val usPersonDecoder: Decoder[Person] = deriveDecoder[Person].emap { case p if p.country == "us" => Right(p) case _ => Left("Person's country is not 'us'") } val personList: List[Person] = Source.fromResource("Persons.json") .getLines .flatMap { line => decode[Person](line)(usPersonDecoder) match { case Right(p) => Some(p) // 解码失败包含格式错误和country不符合,可选择忽略或抛出异常 case Left(ex) => None // case Left(ex) => throw new RuntimeException(ex) } } .toList
关键注意点
- 超大文件必须用
getLines的Iterator流式处理,不要先转成List,否则会把整个文件加载到内存,引发内存溢出 - 解码失败的处理可根据业务需求调整:是抛出异常终止程序,还是忽略错误行继续处理
内容的提问来源于stack exchange,提问作者jbogart
相关产品推荐
相关产品推荐

