使用Apache Rya批量加载数据时遇Could not find or load main class hdfs错误
Let's break down what's happening here and fix this issue step by step:
What's Causing the Error?
Looking at the error logs, Hadoop's shell script is mistakenly interpreting your HDFS jar path (hdfs://localhost:9000/user/rya.mapreduce-4.0.0-incubating-shaded.jar) as part of an environment variable name (like HADOOP_HDFS://LOCALHOST:9000/USER/RYA.MAPREDUCE-4.0.0-INCUBATING-SHADED.JAR_USER). This happens because the hadoop jar command expects a local file system path by default, not an HDFS path. When you pass an HDFS URL directly, the script tries to parse it as a variable prefix, leading to invalid variable name errors. The final "could not find main class" error is a side effect—Hadoop is treating the HDFS path as the main class name instead of recognizing it as the jar file.
Fixes to Try
1. Use a Local Copy of the Shaded Jar (Most Reliable)
The simplest fix is to download the jar from HDFS to your local node first, then run the command with the local path:
# Download the jar from HDFS to a local temp directory bin/hadoop fs -get /user/rya.mapreduce-4.0.0-incubating-shaded.jar /tmp/ # Run the Rya batch load command with the local jar path bin/hadoop jar /tmp/rya.mapreduce-4.0.0-incubating-shaded.jar org.apache.rya.mapreduce.LoadRyaBatchJob [your-job-arguments]
Make sure to replace [your-job-arguments] with the actual parameters required by the LoadRyaBatchJob (like input path, Rya instance details, etc., as specified in the official docs).
2. Correctly Reference the HDFS Jar (If Your Hadoop Version Supports It)
If you prefer to use the HDFS jar directly, ensure you wrap the path in quotes to prevent the shell script from misinterpreting special characters like :// and /:
bin/hadoop jar "hdfs://localhost:9000/user/rya.mapreduce-4.0.0-incubating-shaded.jar" org.apache.rya.mapreduce.LoadRyaBatchJob [your-job-arguments]
Note: Not all Hadoop versions support loading jars directly from HDFS, so this might not work for older distributions. The local copy method is more universally compatible.
3. Double-Check the Main Class Name
The error about the missing main class confirms that Hadoop is misreading your command structure. Ensure you're specifying the correct main class right after the jar path. For Rya's batch load, the main class should be org.apache.rya.mapreduce.LoadRyaBatchJob—don't skip this parameter, and don't mix up the order of the jar path and main class.
Quick Verification Step
Before re-running the command, confirm that the local jar has the main class you need:
jar -tf /tmp/rya.mapreduce-4.0.0-incubating-shaded.jar | grep LoadRyaBatchJob
This should output the full path to the main class, confirming it exists in the jar.
内容的提问来源于stack exchange,提问作者M.Taki_Eddine

