You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark Glue将bigint写入ORC格式的正确性及预期行为咨询

Understanding Spark/AWS Glue Type Mapping for CSV to ORC

Let's walk through exactly what's happening with your code, clear up some misconceptions, and answer your questions about DecimalType and ORC compatibility.

First, Clarify Key Type Mappings

Quick note: I suspect there's a typo in your applyMapping call ("biging_col" should probably be "bigint_col"), but let's focus on the core logic here.

When you use applyMapping to convert a string column to bigint in Glue, here's the under-the-hood mapping:

  • Glue's bigint type directly maps to Spark's LongType — a 64-bit signed integer, equivalent to Java's Long.
  • The conversion parses string values from your CSV into Long instances. If any value falls outside the Long range (less than -9223372036854775808 or greater than 9223372036854775807), Glue will either throw an error or route the bad record to an error table (depending on your job's error handling setup). Since your code runs without issues, your CSV values are all within the valid Long range.

2. ORC Format Type Support

Your assumption that ORC lacks Decimal or bigint support is incorrect:

  • ORC has native support for bigint (64-bit integers), which aligns perfectly with Spark's LongType.
  • ORC also supports decimal types (since ORC 1.1), which maps directly to Spark's DecimalType.

What Actually Happens in Your Workflow?

Let's break down each step:

  • Reading CSV: Glue pulls in biging_col as a string, storing it in a DynamicFrame.
  • Applying Mapping: Valid integer strings are converted to Glue bigint (Spark LongType). Invalid values are handled per your error configuration.
  • Writing ORC: Spark serializes the LongType column into ORC's native bigint type. The ORC file will store these as pure 64-bit integers — no .0 decimal suffix, no precision loss.

Your DecimalType Question

You asked if careful use of DecimalType avoids conversion to double:

  • Absolutely! Spark's DecimalType uses java.math.BigDecimal under the hood, which preserves exact precision. If you define a DecimalType with a scale of 0 (e.g., DecimalType(38, 0)), it will store integers of up to 38 digits without any decimal points or conversion to double.
  • This is the right approach if your CSV values exceed the Long range. You'd adjust your applyMapping to map the string column to decimal(38,0) instead of bigint, and ORC will store this as its native decimal type.

Expected Behavior

  • For values within the Long range: Your workflow works as intended — ORC will have a bigint column with exact integer values, no precision loss, no .0 suffix.
  • For values outside Long range: Switch to DecimalType(precision, 0) to avoid errors and preserve exact values. ORC will handle this decimal type natively.

内容的提问来源于stack exchange,提问作者Cherry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:22:11