You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark Scala处理小数格式:实现8.7位Decimal类型转换

Solution for Converting to Decimal(15,7) (8.7 Format)

The core issue here is that your raw input (like 7373743343333444.) is actually a concatenated string of integer and decimal parts with a trailing dot, not a properly formatted decimal value. Directly casting it to DecimalType(38,3) just treats the entire numeric portion as an integer and pads zeros to the decimal places, which doesn't align with your 8-integer, 7-decimal requirement.

Here are two targeted solutions based on your exact needs:

Option 1: Treat the raw string as a numeric value missing the decimal point (right-shift 7 places)

This assumes your raw data is a full number where the last 7 digits should be the decimal portion, and the rest are the integer portion (capped at 8 digits to fit your 8.7 format):

import org.apache.spark.sql.functions._
import org.apache.spark.sql.types.DecimalType

// Step 1: Clean the input string by removing the trailing dot
val cleanedDF = dataFileWithSchema.withColumn(
  "cleaned_num",
  regexp_replace(col("COLUMN NAME"), "\\.$", "") // Remove trailing '.'
)

// Step 2: Convert to a large decimal, divide by 10^7 to shift decimal 7 places left, then cast to Decimal(15,7)
val resultDF = cleanedDF.withColumn(
  "COLUMN NAME",
  (col("cleaned_num").cast(DecimalType(38, 0)) / lit(10000000))
    .cast(DecimalType(15, 7)) // 15 total digits = 8 integer +7 decimal
)

Note: If your raw numeric string is longer than 15 digits (e.g., 16 digits), this will result in an integer portion longer than 8 digits, which will cause an overflow when casting to Decimal(15,7). You'll need to decide whether to truncate, round, or reject these values if they exist in your data.

Option 2: Explicitly split into 8-digit integer and 7-digit decimal parts

If you need to strictly take the first 8 digits as the integer portion and the next 7 digits as the decimal portion (ignoring any extra digits), use this approach:

import org.apache.spark.sql.functions._
import org.apache.spark.sql.types.DecimalType

// Step 1: Clean the input string
val cleanedDF = dataFileWithSchema.withColumn(
  "cleaned_num",
  regexp_replace(col("COLUMN NAME"), "\\.$", "")
)

// Step 2: Split into integer and decimal parts, pad decimal part to 7 digits if needed
val splitDF = cleanedDF.withColumn(
  "integer_part", substring(col("cleaned_num"), 1, 8) // First 8 digits
).withColumn(
  "decimal_part", lpad(substring(col("cleaned_num"), 9, 7), 7, "0") // Next 7 digits, pad with zeros if shorter
)

// Step 3: Combine and cast to target Decimal type
val resultDF = splitDF.withColumn(
  "COLUMN NAME",
  concat(col("integer_part"), lit("."), col("decimal_part"))
    .cast(DecimalType(15, 7))
)

This will always produce a value with exactly 8 integer digits and 7 decimal digits, regardless of the original string length (extra digits are truncated, shorter strings are padded with zeros).

You can verify the result with resultDF.show(5) to confirm it matches your 8.7 format requirement.


内容的提问来源于stack exchange,提问作者Venkat J

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:42:43