使用Bulk Load向HBase导入海量数据遇log4j权限错误的解决方案
Fixing Log4j Permission Issues
The errors you're encountering stem from the user executing HBase commands not having write access to the /var/log/hbase directory. Here are actionable fixes:
1. Run commands as the hbase user
Use sudo to execute commands under the hbase user, which has native permissions for HBase log directories:
sudo -u hbase hbase org.apache.hadoop.hbase.mapreduce.ImportTsv [your-import-arguments] sudo -u hbase hbase org.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles /storefile flat_table_sf10
2. Override log4j configuration temporarily
Specify a custom log directory you control when running commands, bypassing the system /var/log/hbase path:
hbase org.apache.hadoop.hbase.mapreduce.ImportTsv \ -Dlog4j.configuration=file:/home/your-user/custom-log4j.properties \ [your-import-arguments]
Create custom-log4j.properties with writable paths:
log4j.rootLogger=INFO,DRFA log4j.appender.DRFA=org.apache.log4j.DailyRollingFileAppender log4j.appender.DRFA.File=/home/your-user/hbase-logs/hbase.log log4j.appender.DRFA.DatePattern='.'yyyy-MM-dd log4j.appender.DRFA.layout=org.apache.log4j.PatternLayout log4j.appender.DRFA.layout.ConversionPattern=%d{ISO8601} %p %c: %m%n log4j.logger.SecurityLogger=INFO,DRFAS log4j.appender.DRFAS=org.apache.log4j.DailyRollingFileAppender log4j.appender.DRFAS.File=/home/your-user/hbase-logs/SecurityAuth.audit log4j.appender.DRFAS.DatePattern='.'yyyy-MM-dd log4j.appender.DRFAS.layout=org.apache.log4j.PatternLayout log4j.appender.DRFAS.layout.ConversionPattern=%d{ISO8601} %p %c: %m%n
3. Adjust directory permissions (non-production only)
Grant temporary write access to your user on the system log directory (not recommended for production due to security risks):
sudo chmod o+w /var/log/hbase
More Efficient Bulk Load Methods for AWS EMR
For 50 million rows, these optimizations will speed up your workflow:
1. Use Spark instead of MapReduce
Spark offers faster parallel processing for large datasets. Use the Spark-HBase connector to generate StoreFiles directly:
import org.apache.hadoop.hbase.spark.HBaseContext import org.apache.hadoop.hbase.{HBaseConfiguration, TableName} import org.apache.hadoop.hbase.client.Put import org.apache.hadoop.hbase.io.ImmutableBytesWritable import org.apache.hadoop.hbase.util.Bytes val conf = HBaseConfiguration.create() val hbaseContext = new HBaseContext(spark.sparkContext, conf) // Load TSV data from HDFS/S3 val tsvData = spark.read.option("delimiter", "\t").csv("hdfs://path/to/tsv-files").rdd // Convert rows to HBase Put objects val hbasePuts = tsvData.map(row => { val put = new Put(Bytes.toBytes(row.getString(0))) // Row key from first column put.addColumn(Bytes.toBytes("cf"), Bytes.toBytes("col1"), Bytes.toBytes(row.getString(1))) // Add other columns as needed (new ImmutableBytesWritable, put) }) // Bulk load into HBase hbaseContext.bulkLoad(hbasePuts, TableName.valueOf("flat_table_sf10"))
2. Skip HDFS copy - load directly from S3
Eliminate the HDFS transfer step by using HBase/S3 native integration:
hbase org.apache.hadoop.hbase.mapreduce.ImportTsv \ -Dimporttsv.bulk.output=s3://your-bucket/storefile-output \ -Dimporttsv.columns=HBASE_ROW_KEY,cf:col1,cf:col2 \ flat_table_sf10 \ s3://your-bucket/input-tsv-files/*.tsv
3. Optimize EMR Cluster Settings
- Use compute-optimized EC2 instances to handle CPU-intensive bulk load tasks
- Increase parallelism with mapper/reducer tuning:
-Dmapreduce.job.maps=64 -Dmapreduce.job.reduces=32 - Enable HBase optimizations like short-circuit reads and off-heap caching via EMR configuration
内容的提问来源于stack exchange,提问作者Lucas Emanuel

