Scala代码报ArrayIndexOutOfBoundsException:5,求XML数据长度计算及第7列解析方案
解决Scala代码中的ArrayIndexOutOfBoundsException并实现需求
Hey there, let's break down why you're hitting that java.lang.ArrayIndexOutOfBoundsException: 5 error and adjust the code to meet your two requirements: calculating the length of XML data in your delimited file, and parsing the 7th column.
Why the Error Happens
The core issue here is the order of operations in your code:
- You're first running a
mapthat tries to accessevent(5)(the 6th column, since arrays are 0-indexed) before checking if the line actually has enough columns. - The
filterthat checksline._4.length==7comes after thismap, so by the time you get to filtering, the code has already tried to access an index that doesn't exist in shorter lines—hence the out-of-bounds exception.
Fixed Code to Meet Your Requirements
Let's restructure the code to first filter invalid lines, then handle your XML length calculation and 7th column parsing. Here's a revised version with explanations:
import scala.xml._ // Import for XML parsing // Assuming `data` is your input RDD/Dataset of lines from the delimited file val result = data // Split each line by the ^ delimiter first .map(line => line.split("\\^")) // FIRST: Filter out lines that don't have exactly 7 columns (avoids index errors) .filter(arr => arr.length == 7) // Now process valid lines: extract needed values, calculate XML length, parse XML .map(arr => { val eventType = Integer.parseInt(arr(2)) // Your original `x` value val sixthColumn = arr(5) // Your original `y` value (used in filtering) val seventhColumn = arr(6) // The 7th column (XML data, 0-indexed) val xmlLength = seventhColumn.length // Requirement 1: Calculate XML data length // Requirement 2: Parse the 7th column (XML) // Add error handling in case of malformed XML val parsedXmlRoot = try { XML.loadString(seventhColumn).label // Example: get root element tag } catch { case e: Exception => "Malformed XML" } // Bundle values into a tuple for filtering and output (eventType, sixthColumn, seventhColumn, xmlLength, parsedXmlRoot) }) // Apply your original filtering conditions .filter(tuple => tuple._1 == 100 && tuple._2.nonEmpty) // Adjust output to whatever you need (here we keep key values) .map(tuple => (tuple._1, tuple._3, tuple._4, tuple._5)) // Print results println(s"Number of valid records: ${result.count()}") result.foreach(println)
Key Fixes & Improvements
- Moved filtering first: We now check for exactly 7 columns before trying to access any indexes, eliminating the out-of-bounds error entirely.
- Explicit 7th column access: The 7th column is at index
6(since arrays start at 0), so we usearr(6)to target it directly. - XML handling: Added XML parsing with error handling to avoid crashes from malformed XML, and included the length calculation as required.
- Clearer structure: Broke down operations into logical steps so it's easier to adjust later if your requirements change.
内容的提问来源于stack exchange,提问作者Mahesh Yadav
相关产品推荐
相关产品推荐

