关于Tom White《Hadoop》第4版JAR文件用途及自制JAR使用的技术问询
Answers to Your Hadoop JAR & Custom Built JAR Questions
1. JAR File Purposes from Tom White's Hadoop: The Definitive Guide, 4th Edition
Tom White's book breaks down Hadoop's ecosystem into core components, and the JARs mentioned are the foundational building blocks for each subsystem. Here's a clear breakdown of the key ones and their roles:
hadoop-common.jar: This is the backbone JAR shared across all Hadoop modules. It contains universal utilities, configuration classes (likeConfiguration), I/O abstractions, and common data structures that glue HDFS, MapReduce, and YARN together. The book emphasizes this as the starting point for any Hadoop development.hadoop-hdfs.jar: This JAR powers the Hadoop Distributed File System (HDFS). It includes code for the NameNode, DataNode, client APIs for interacting with HDFS (likeFileSystem), and all logic for block storage, replication, and metadata management—critical for the distributed storage layer covered in the book's HDFS chapters.hadoop-mapreduce-client-core.jar: For MapReduce development, this JAR holds the core classes you’ll work with:Mapper,Reducer,Job,JobConf, and runtime logic for task execution. The book relies heavily on this JAR when walking through custom MapReduce job creation, as it’s required to compile and run these applications.hadoop-yarn-common.jar&hadoop-yarn-server-resourcemanager.jar: These are central to YARN (Yet Another Resource Negotiator), Hadoop’s cluster resource management layer. The 4th edition covers YARN’s role in allocating cluster resources, and these JARs contain the ResourceManager, NodeManager, and client classes for submitting applications to the cluster.hadoop-client.jar: A convenience JAR that bundles client-side classes from HDFS, MapReduce, and YARN. The book mentions this to simplify development—you can include one JAR instead of multiple component-specific ones when building client applications.
2. Figuring Out Usage for a JAR Built from a Git Repository
If you’ve built a JAR from a cloned Git repo but aren’t sure what it does or how to use it, follow these practical, actionable steps:
- Start with the repository’s docs: Look for a
README.md,docs/folder, orUSAGEfile. Most projects explicitly outline the JAR’s purpose (is it a command-line tool, a library, or a service?), required dependencies, and basic usage examples here—this is the fastest path to clarity. - Inspect the JAR’s contents: Run
jar tf your-jar-file.jarto list all files inside. Key things to look for:META-INF/MANIFEST.MF: Check for aMain-Classentry—if present, this means the JAR is executable. Try running it withjava -jar your-jar-file.jar(you might need additional dependencies in the classpath, which should be noted in the manifest or docs).- Package names and class files: These often hint at functionality (e.g.,
com.example.data.pipelinesuggests a data processing tool).
- Test execution (if applicable): If there’s a
Main-Class, run the JAR with the--helpor-hflag—most executable JARs include built-in usage instructions. - Check build files: Look for
pom.xml(Maven) orbuild.gradle(Gradle) in the repo. These files list dependencies the JAR needs, and also show how it’s intended to be used (e.g., as a dependency in another project via Maven/Gradle coordinates). - Dig into repo history/issues: If docs are sparse, check commit messages or open/closed issues. Other users might have asked similar questions, or commits might mention the JAR’s intended use case.
内容的提问来源于stack exchange,提问作者Abhishek Misra
相关产品推荐
相关产品推荐

