You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Elasticsearch生成随机测试记录?需百万级记录验证内存

如何为Elasticsearch生成百万级测试记录

嘿,要生成100万+测试记录来验证Elasticsearch的内存占用,直接循环逐个插入可不行——不仅慢到离谱,还容易踩坑。我来给你梳理下正确的姿势,顺便优化你的原有代码:

先修正原代码的核心问题

你的代码里有两个致命问题,必须先解决:

  • ID重复覆盖:id: '1'会导致所有新插入的记录都覆盖同一条数据,最后你只能得到1条记录,完全达不到测试目的
  • 单条插入效率极低:逐个调用client.index()会产生上百万次HTTP请求,ES的性能会被这种高频小请求拖垮,而且你的客户端也可能因为请求堆积崩溃

优化后的批量插入方案(Node.js示例)

Elasticsearch提供了Bulk API,专门用来批量处理文档写入,这才是生成百万级数据的正确方式。下面是完整的代码示例,基于官方的@elastic/elasticsearch客户端:

const { Client } = require('@elastic/elasticsearch');
const client = new Client({ node: 'http://localhost:9200' });

// 提前创建测试索引,调整参数提升写入性能
async function createTestIndex() {
  try {
    await client.indices.create({
      index: 'test',
      body: {
        settings: {
          number_of_shards: 3, // 根据你的节点数调整,建议每个分片不超过50GB
          number_of_replicas: 0, // 插入期间关闭副本,完成后再开启
          refresh_interval: '-1' // 插入期间不自动刷新,大幅提升写入速度
        },
        mappings: {
          properties: {
            timestamp: { type: 'date' },
            random_field: { type: 'keyword' },
            numeric_field: { type: 'integer' }
          }
        }
      }
    });
    console.log('测试索引创建成功');
  } catch (err) {
    if (err.meta.statusCode === 400) {
      console.log('索引已存在,跳过创建');
    } else {
      throw err;
    }
  }
}

// 生成批量插入的请求体
function generateBulkBatch(startId, batchSize) {
  const bulkBody = [];
  const now = Date.now();
  
  for (let i = 0; i < batchSize; i++) {
    const docId = startId + i;
    // 生成模拟的测试数据,你可以根据需求调整字段
    const timestamp = new Date(now - Math.floor(Math.random() * 30 * 24 * 60 * 60 * 1000)); // 近30天的随机时间
    const randomValue = `test_value_${Math.floor(Math.random() * 1000000)}`;
    const numericValue = Math.floor(Math.random() * 1000);

    bulkBody.push({
      index: { _index: 'test', _id: docId.toString() }
    });
    bulkBody.push({
      timestamp: timestamp.toISOString(),
      random_field: randomValue,
      numeric_field: numericValue
    });
  }
  return bulkBody;
}

// 批量插入百万条数据
async function insertMillionRecords(totalCount = 1000000, batchSize = 1000) {
  await createTestIndex();
  
  const totalBatches = Math.ceil(totalCount / batchSize);
  let successCount = 0;

  for (let batch = 0; batch < totalBatches; batch++) {
    const startId = batch * batchSize;
    const currentBatchSize = batch === totalBatches - 1 ? totalCount - startId : batchSize;
    const bulkBody = generateBulkBatch(startId, currentBatchSize);

    try {
      const response = await client.bulk({ body: bulkBody });
      
      // 统计成功插入的数量
      response.items.forEach(item => {
        if (item.index.status === 201) successCount++;
      });

      console.log(`完成第${batch+1}/${totalBatches}批插入,累计成功${successCount}条`);
    } catch (err) {
      console.error(`第${batch+1}批插入失败:`, err.message);
    }
  }

  // 插入完成后,恢复索引的自动刷新和副本设置
  await client.indices.putSettings({
    index: 'test',
    body: {
      refresh_interval: '1s',
      number_of_replicas: 1 // 根据你的需求调整副本数
    }
  });

  console.log(`百万条数据插入完成,最终成功${successCount}条`);
}

// 执行插入
insertMillionRecords().catch(err => console.error('整体执行失败:', err));

关键优化点说明

  • 批量提交:每1000条数据提交一次,平衡请求大小和性能,你可以根据ES的配置调整batchSize(建议500-2000之间)
  • 索引参数优化:插入期间关闭副本、禁用自动刷新,能把写入性能提升数倍
  • 模拟真实数据:生成包含日期、字符串、数值的多字段文档,更贴近真实业务场景,测试内存占用更准确
  • 唯一ID:用自增ID确保每条记录都能被正确保存

内存测试的额外建议

  • 插入完成后,执行一次手动刷新:await client.indices.refresh({ index: 'test' }),让所有数据都被加载到内存
  • 查看内存占用可以用ES的命令:
    • 查看节点内存:curl http://localhost:9200/_cat/nodes?v(关注heap.percent和heap.max字段)
    • 查看索引内存占用:curl http://localhost:9200/_cat/indices?v(关注pri.store.size和store.size字段)
  • 如果要测试极端内存情况,可以增加数组、嵌套对象等复杂字段,或者调整字段的分词方式(比如text字段会占用更多内存)

内容的提问来源于stack exchange,提问作者Palaniichuk Dmytro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:22:44