You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark ML:如何将Scala计数字典转换为保留频率的机器学习特征?

How to Convert Location Frequency Maps to ML Features (Preserving Counts)

Great question! You’re absolutely right that using CountVectorizer on a string list throws away the explicit frequency data you already have, which is a waste since those counts are meaningful features. Let’s go through practical, Scala-based approaches (focused on Spark ML since you mentioned CountVectorizer) to preserve that information:


1. Explicit Vector Construction (Dense/Sparse)

If you’re working with in-memory data or small datasets, you can directly convert each map into a numerical vector where each index corresponds to a unique location, and the value is its count.

Step 1: Define your feature space

First, collect all unique locations across all your maps to establish a consistent feature order:

import org.apache.spark.ml.linalg.{Vectors, Vector}

// Sample input data
val locationMaps = Seq(
  Map("beach" -> 31, "cafe" -> 140, "prison" -> 2),
  Map("beach" -> 5, "park" -> 20)
)

// Get sorted list of all unique locations
val allLocations = locationMaps.flatMap(_.keys).distinct.sorted
// Map each location to its index in the feature vector
val locationToIndex = allLocations.zipWithIndex.toMap

Step 2: Convert maps to vectors

Choose between dense vectors (good for small feature spaces) or sparse vectors (more memory-efficient for large spaces with many zero values):

// Dense vector: includes all locations, even those with 0 counts
def mapToDenseVector(locationMap: Map[String, Int]): Vector = {
  val values = allLocations.map(loc => locationMap.getOrElse(loc, 0).toDouble).toArray
  Vectors.dense(values)
}

// Sparse vector: only stores non-zero counts and their indices
def mapToSparseVector(locationMap: Map[String, Int]): Vector = {
  val (indices, values) = locationMap.toSeq
    .filter { case (_, count) => count > 0 }
    .map { case (loc, count) => (locationToIndex(loc), count.toDouble) }
    .unzip
  Vectors.sparse(allLocations.size, indices.toArray, values.toArray)
}

// Apply to all maps
val featureVectors = locationMaps.map(mapToDenseVector)

2. Spark DataFrame + VectorAssembler (Scalable for Big Data)

For large-scale datasets, use Spark’s DataFrame API to pivot your data into location-specific columns, then assemble them into a feature vector. This integrates seamlessly with Spark ML pipelines.

Step 1: Convert maps to a long-format DataFrame

import org.apache.spark.sql.SparkSession
import org.apache.spark.ml.feature.VectorAssembler

val spark = SparkSession.builder().appName("LocationFeatures").getOrCreate()
import spark.implicits._

// Convert maps to (id, location, count) rows
val df = locationMaps.zipWithIndex.flatMap { case (map, id) =>
  map.map { case (loc, cnt) => (id, loc, cnt) }
}.toDF("id", "location", "count")

Step 2: Pivot to wide format and assemble features

// Pivot to get each location as a column with its count (fill missing with 0)
val pivotedDf = df.groupBy("id").pivot("location").sum("count").na.fill(0)

// Get list of location columns (exclude the ID column)
val locationCols = pivotedDf.columns.filter(_ != "id")

// Assemble columns into a single feature vector
val assembler = new VectorAssembler()
  .setInputCols(locationCols)
  .setOutputCol("features")

val featureDf = assembler.transform(pivotedDf)

Key Notes

  • Avoid CountVectorizer here: CountVectorizer is designed to count tokens from raw text. Using it would require exploding your map into a repetitive string list (e.g., ["beach"] * 31 + ["cafe"] * 140), which is inefficient and redundant since you already have precomputed counts.
  • Normalization (optional): If your ML model benefits from scaled features, apply StandardScaler or MinMaxScaler to the resulting vector to normalize counts to a standard range.
  • Handling new locations: For future data with unseen locations, decide whether to ignore them, add them to your feature space (and retrain), or map them to an "unknown" category.

内容的提问来源于stack exchange,提问作者Aivaras

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:09:37