Spark Glue将bigint写入ORC格式的正确性及预期行为咨询
Understanding Spark/AWS Glue Type Mapping for CSV to ORC
Let's walk through exactly what's happening with your code, clear up some misconceptions, and answer your questions about DecimalType and ORC compatibility.
First, Clarify Key Type Mappings
Quick note: I suspect there's a typo in your applyMapping call ("biging_col" should probably be "bigint_col"), but let's focus on the core logic here.
1. Glue DynamicFrame ↔ Spark DataFrame Type Links
When you use applyMapping to convert a string column to bigint in Glue, here's the under-the-hood mapping:
- Glue's
biginttype directly maps to Spark'sLongType— a 64-bit signed integer, equivalent to Java'sLong. - The conversion parses string values from your CSV into
Longinstances. If any value falls outside theLongrange (less than-9223372036854775808or greater than9223372036854775807), Glue will either throw an error or route the bad record to an error table (depending on your job's error handling setup). Since your code runs without issues, your CSV values are all within the validLongrange.
2. ORC Format Type Support
Your assumption that ORC lacks Decimal or bigint support is incorrect:
- ORC has native support for
bigint(64-bit integers), which aligns perfectly with Spark'sLongType. - ORC also supports
decimaltypes (since ORC 1.1), which maps directly to Spark'sDecimalType.
What Actually Happens in Your Workflow?
Let's break down each step:
- Reading CSV: Glue pulls in
biging_colas a string, storing it in a DynamicFrame. - Applying Mapping: Valid integer strings are converted to Glue
bigint(SparkLongType). Invalid values are handled per your error configuration. - Writing ORC: Spark serializes the
LongTypecolumn into ORC's nativebiginttype. The ORC file will store these as pure 64-bit integers — no.0decimal suffix, no precision loss.
Your DecimalType Question
You asked if careful use of DecimalType avoids conversion to double:
- Absolutely! Spark's
DecimalTypeusesjava.math.BigDecimalunder the hood, which preserves exact precision. If you define aDecimalTypewith a scale of0(e.g.,DecimalType(38, 0)), it will store integers of up to 38 digits without any decimal points or conversion todouble. - This is the right approach if your CSV values exceed the
Longrange. You'd adjust yourapplyMappingto map the string column todecimal(38,0)instead ofbigint, and ORC will store this as its native decimal type.
Expected Behavior
- For values within the
Longrange: Your workflow works as intended — ORC will have abigintcolumn with exact integer values, no precision loss, no.0suffix. - For values outside
Longrange: Switch toDecimalType(precision, 0)to avoid errors and preserve exact values. ORC will handle this decimal type natively.
内容的提问来源于stack exchange,提问作者Cherry
相关产品推荐
相关产品推荐

