使用Order by/Group by查询Hive表时出现执行错误求助
Problem Description
I'm using DBVisualizer on Mac to query Hive tables, with the Simba Hive JDBC41 driver (version 1.0.42.1054) installed. Simple queries like select * from table_name limit 10 work fine, but queries with ORDER BY or GROUP BY (e.g., select * from table_name order by rollingtime limit 10) fail with the following error:
09:56:17 START Executing for: 'NewDev' [Hive], Database: Hive, Schema: sdc
09:56:17 FAILED [SELECT - 0 rows, 0.504 secs] [Code: 500051, SQL State: HY000] [Simba]HiveJDBCDriver ERROR processing query/statement. Error Code: 2, SQL state: Error while processing statement: FAILED: Execution Error, return code 2 from org.apache.hadoop.hive.ql.exec.tez.TezTask. Vertex failed, vertexName=Map 1, vertexId=vertex_1516123265840_0008_8_00, diagnostics=[Task failed, taskId=task_1516123265840_0008_8_00_000000, diagnostics=[TaskAttempt 0 failed, info=[Error: Error while running task ( failure ) : java.lang.NoClassDefFoundError: Could not initialize class org.apache.tez.runtime.library.api.TezRuntimeConfiguration at org.apache.tez.runtime.library.output.OrderedPartitionedKVOutput.start(OrderedPartitionedKVOutput.java:107) at org.apache.hadoop.hive.ql.exec.tez.MapRecordProcessor.init(MapRecordProcessor.java:186) at org.apache.hadoop.hive.ql.exec.tez.TezProcessor.initializeAndRunProcessor(TezProcessor.java:188) at org.apache.hadoop.hive.ql.exec.tez.TezProcessor.run(TezProcessor.java:172) at org.apache.tez.runtime.LogicalIOProcessorRuntimeTask.run(LogicalIOProcessorRuntimeTask.java:370) at org.apache.tez.runtime.task.TaskRunner2Callable$1.run(TaskRunner2Callable.java:73) at org.apache.tez.runtime.task.TaskRunner2Callable$1.run(TaskRunner2Callable.java:61) at java.security.AccessController.doPrivileged(Native Method) at javax.security.auth.Subject.doAs(Subject.java:422) at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1866) at org.apache.tez.runtime.task.TaskRunner2Callable.callInternal(TaskRunner2Callable.java:61) at org.apache.tez.runtime.task.TaskRunner2Callable.callInternal(TaskRunner2Callable.java:37) at org.apache.tez.common.CallableWithNdc.call(CallableWithNdc.java:36) at org.apache.hadoop.hive.llap.daemon.impl.StatsRecordingThreadPool$WrappedCallable.call(StatsRecordingThreadPool.java:110) at java.util.concurrent.FutureTask.run(FutureTask.java:266) at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) at java.lang.Thread.run(Thread.java:745) , errorMessage=Cannot recover from this error:java.lang.NoClassDefFoundError: Could not initialize class org.apache.tez.runtime.library.api.TezRuntimeConfiguration at org.apache.tez.runtime.library.output.OrderedPartitionedKVOutput.start(OrderedPartitionedKVOutput.java:107) at org.apache.hadoop.hive.ql.exec.tez.MapRecordProcessor.init(MapRecordProcessor.java:186) at org.apache.hadoop.hive.ql.exec.tez.TezProcessor.initializeAndRunProcessor(TezProcessor.java:188) at org.apache.hadoop.hive.ql.exec.tez.TezProcessor.run(TezProcessor.java:172) at org.apache.tez.runtime.LogicalIOProcessorRuntimeTask.run(LogicalIOProcessorRuntimeTask.java:370) at org.apache.tez.runtime.task.TaskRunner2Callable$1.run(TaskRunner2Callable.java:73) at org.apache.tez.runtime.task.TaskRunner2Callable$1.run(TaskRunner2Callable.java:61) at java.security.AccessController.doPrivileged(Native Method) at javax.security.auth.Subject.doAs(Subject.java:422) at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1866) at org.apache.tez.runtime.task.TaskRunner2Callable.callInternal(TaskRunner2Callable.java:61) at org.apache.tez.runtime.task.TaskRunner2Callable.callInternal(TaskRunner2Callable.java:37) at org.apache.tez.common.CallableWithNdc.call(CallableWithNdc.java:36) at org.apache.hadoop.hive.llap.daemon.impl.StatsRecordingThreadPool$WrappedCallable.call(StatsRecordingThreadPool.java:110) at java.util.concurrent.FutureTask.run(FutureTask.java:266) at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142) at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617) at java.lang.Thread.run(Thread.java:745) ]], Vertex did not succeed due to OWN_TASK_FAILURE, failedTasks:1 killedTasks:1, Vertex vertex_1516123265840_0008_8_00 [Map 1] killed/failed due to:OWN_TASK_FAILURE]Vertex killed, vertexName=Reducer 2, vertexId=vertex_1516123265840_0008_8_01, diagnostics=[Vertex received Kill while in RUNNING state., Vertex did not succeed due to OTHER_VERTEX_FAILURE, failedTasks:0 killedTasks:1, Vertex vertex_1516123265840_0008_8_01 [Reducer 2] killed/failed due to:OTHER_VERTEX_FAILURE]DAG did not succeed due to VERTEX_FAILURE. failedVertices:1 killedVertices:1, Query: select * from nomura_qa_mblock_capacity_stage order by rollingtime limit 10. select * from nomura_qa_mblock_capacity_stage order by rollingtime limit 10;
09:56:17 END Execution 1 statement(s) executed, 0 row(s) affected, exec/fetch time: 0.504/0.000 secs [0 successful, 1 errors]
Solution
I’ve dealt with this exact issue before when working with Simba JDBC drivers and Hive’s Tez execution engine. Here’s how to fix it:
First, Understand the Root Cause
The NoClassDefFoundError for TezRuntimeConfiguration means your JDBC driver either lacks the necessary Tez runtime libraries, or the driver version is incompatible with the Tez version running on your Hive cluster. Simple queries don’t trigger Tez (they often run on MapReduce or LLAP without needing the full Tez stack), which is why they work without issues.
Fix Options
1. Quick Workaround: Force MapReduce Execution
If you don’t need Tez for your workloads, you can override the execution engine directly in DBVisualizer:
- Open your Hive connection in DBVisualizer
- Navigate to the Properties tab
- Add a new property:
hive.execution.enginewith the valuemr - Save the connection and retry your
ORDER BY/GROUP BYquery
This skips Tez entirely, avoiding the missing class error.
2. Update to a Compatible JDBC Driver Version
The 1.0.42 driver you’re using is likely outdated and missing newer Tez dependencies. To fix this:
- First, check your HDP cluster’s Hive and Tez versions (match the driver to your cluster’s release)
- Download the corresponding Simba Hive JDBC driver for your HDP version
- Replace the old driver JARs in DBVisualizer’s driver directory with the new ones
- Restart DBVisualizer and test your queries again
3. Manually Add Missing Tez Dependencies
If updating the driver isn’t feasible, you can add the required Tez JARs to DBVisualizer’s classpath:
- Grab the
tez-runtime-library-*.jarand related Tez configuration JARs from your HDP cluster (you can copy them from the cluster’s Tez installation directory) - In DBVisualizer, go to Tools > Driver Manager
- Select your Hive driver and click Edit
- Go to the Files tab, add the missing Tez JARs, and save
- Restart DBVisualizer and re-run your queries
Additional Checks
- Run the failing query directly on the cluster using Hive CLI to confirm the issue isn’t on the cluster side
- Ensure DBVisualizer is using a Java version compatible with your Hive/Tez stack (Java 8 is typically required for older HDP releases)
内容的提问来源于stack exchange,提问作者user4829568

