Spark ML:如何将Scala计数字典转换为保留频率的机器学习特征?
Great question! You’re absolutely right that using CountVectorizer on a string list throws away the explicit frequency data you already have, which is a waste since those counts are meaningful features. Let’s go through practical, Scala-based approaches (focused on Spark ML since you mentioned CountVectorizer) to preserve that information:
1. Explicit Vector Construction (Dense/Sparse)
If you’re working with in-memory data or small datasets, you can directly convert each map into a numerical vector where each index corresponds to a unique location, and the value is its count.
Step 1: Define your feature space
First, collect all unique locations across all your maps to establish a consistent feature order:
import org.apache.spark.ml.linalg.{Vectors, Vector} // Sample input data val locationMaps = Seq( Map("beach" -> 31, "cafe" -> 140, "prison" -> 2), Map("beach" -> 5, "park" -> 20) ) // Get sorted list of all unique locations val allLocations = locationMaps.flatMap(_.keys).distinct.sorted // Map each location to its index in the feature vector val locationToIndex = allLocations.zipWithIndex.toMap
Step 2: Convert maps to vectors
Choose between dense vectors (good for small feature spaces) or sparse vectors (more memory-efficient for large spaces with many zero values):
// Dense vector: includes all locations, even those with 0 counts def mapToDenseVector(locationMap: Map[String, Int]): Vector = { val values = allLocations.map(loc => locationMap.getOrElse(loc, 0).toDouble).toArray Vectors.dense(values) } // Sparse vector: only stores non-zero counts and their indices def mapToSparseVector(locationMap: Map[String, Int]): Vector = { val (indices, values) = locationMap.toSeq .filter { case (_, count) => count > 0 } .map { case (loc, count) => (locationToIndex(loc), count.toDouble) } .unzip Vectors.sparse(allLocations.size, indices.toArray, values.toArray) } // Apply to all maps val featureVectors = locationMaps.map(mapToDenseVector)
2. Spark DataFrame + VectorAssembler (Scalable for Big Data)
For large-scale datasets, use Spark’s DataFrame API to pivot your data into location-specific columns, then assemble them into a feature vector. This integrates seamlessly with Spark ML pipelines.
Step 1: Convert maps to a long-format DataFrame
import org.apache.spark.sql.SparkSession import org.apache.spark.ml.feature.VectorAssembler val spark = SparkSession.builder().appName("LocationFeatures").getOrCreate() import spark.implicits._ // Convert maps to (id, location, count) rows val df = locationMaps.zipWithIndex.flatMap { case (map, id) => map.map { case (loc, cnt) => (id, loc, cnt) } }.toDF("id", "location", "count")
Step 2: Pivot to wide format and assemble features
// Pivot to get each location as a column with its count (fill missing with 0) val pivotedDf = df.groupBy("id").pivot("location").sum("count").na.fill(0) // Get list of location columns (exclude the ID column) val locationCols = pivotedDf.columns.filter(_ != "id") // Assemble columns into a single feature vector val assembler = new VectorAssembler() .setInputCols(locationCols) .setOutputCol("features") val featureDf = assembler.transform(pivotedDf)
Key Notes
- Avoid CountVectorizer here: CountVectorizer is designed to count tokens from raw text. Using it would require exploding your map into a repetitive string list (e.g.,
["beach"] * 31 + ["cafe"] * 140), which is inefficient and redundant since you already have precomputed counts. - Normalization (optional): If your ML model benefits from scaled features, apply
StandardScalerorMinMaxScalerto the resulting vector to normalize counts to a standard range. - Handling new locations: For future data with unseen locations, decide whether to ignore them, add them to your feature space (and retrain), or map them to an "unknown" category.
内容的提问来源于stack exchange,提问作者Aivaras

