You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为YCSB基准测试在HBase中创建5GB的测试表?

How to Create a ~5GB Sample Table with HBase & YCSB (Batch Write + Size Control)

Hey there! Since you're new to using HBase with YCSB for benchmarking, let me walk you through exactly how to batch write data and get that table to sit right around 5GB. I’ve messed around with this setup plenty of times, so I’ll share the practical steps that work.

First: Prep Your Environment

Before diving in, make sure:

  • Your HBase cluster is up and running, and you’ve got the hbase-site.xml config file copied into YCSB’s conf directory (this lets YCSB talk to HBase).
  • You’ve created your target table in HBase Shell first. For example:
    create 'ycsb_sample', 'cf'  # 'ycsb_sample' is your table name, 'cf' is the column family
    

Step 1: Configure YCSB for Batch Writing

YCSB’s workload files are where you control how data gets loaded. You can use a pre-built workload (like workloads/workloada) or create a custom one. Here’s what to tweak:

  1. Open your chosen workload file and set these key parameters:

    • recordcount: Total number of rows you want to write (we’ll calculate this later for 5GB)
    • operationcount: Set this equal to recordcount if you’re only doing a load (no read operations yet)
    • fieldcount: Number of columns per row (e.g., 1 for simplicity)
    • fieldlength: Size (in bytes) of each column’s value (this directly impacts row size)
    • requestdistribution: uniform (to spread data evenly across regions)
  2. Run the YCSB load command to start batch writing:

    ./bin/ycsb load hbase12 -P workloads/workloada -p table=ycsb_sample -p columnfamily=cf -p threads=10 -p batchsize=1000
    
    • threads=10: Adjust based on your cluster’s CPU/memory (more threads = faster load, but don’t overload)
    • batchsize=1000: Groups writes into batches to reduce RPC overhead and speed up loading

Step 2: Calculate Record Count to Hit ~5GB

HBase adds small overhead per row (row key, column family name, timestamp, etc.)—plan for roughly 10-20% extra on top of your raw data size. Here’s the simple math:

First, define your row size:
Let’s say you set fieldcount=1 and fieldlength=1024 (1KB of raw data per row). With overhead, each row will take ~1.1KB of storage.

Then calculate total records needed:

Total target size = 5GB = 5 * 1024 * 1024 KB = 5242880 KB
Approx records = 5242880 KB / 1.1 KB per row ≈ 476,000 rows

Round up to recordcount=480000 to account for any variance.

If you want larger rows (e.g., 4KB raw data per row):

fieldlength=4096 → raw row size=4KB → with overhead ~4.4KB per row
Records needed = 5242880 / 4.4 ≈ 119,000 → set recordcount=120000

Verify the Final Size

After loading, check the actual table size:

  • Use HBase Shell to count rows: count 'ycsb_sample'
  • Check HDFS storage usage (HBase stores data here):
    hdfs dfs -du -h /hbase/data/default/ycsb_sample
    

If it’s over/under 5GB, adjust recordcount and re-run the load (you can truncate the table first with truncate 'ycsb_sample').

Pro Tips for Smoother Loading

  • Tweak HBase’s write settings in hbase-site.xml to speed things up:
    • hbase.client.write.buffer: Increase to 10MB (10485760) to reduce RPC calls
    • hbase.hregion.max.filesize: Set to 1GB (1073741824) to avoid frequent region splits during load
  • Start with lower thread counts if you’re testing on a small cluster—too many threads can cause timeouts.
  • If you need consistent row sizes, avoid variable-length fields (stick to fixed fieldlength values).

内容的提问来源于stack exchange,提问作者JrZ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:25:18