如何为Elasticsearch生成随机测试记录?需百万级记录验证内存
如何为Elasticsearch生成百万级测试记录
嘿,要生成100万+测试记录来验证Elasticsearch的内存占用,直接循环逐个插入可不行——不仅慢到离谱,还容易踩坑。我来给你梳理下正确的姿势,顺便优化你的原有代码:
先修正原代码的核心问题
你的代码里有两个致命问题,必须先解决:
- ID重复覆盖:
id: '1'会导致所有新插入的记录都覆盖同一条数据,最后你只能得到1条记录,完全达不到测试目的 - 单条插入效率极低:逐个调用
client.index()会产生上百万次HTTP请求,ES的性能会被这种高频小请求拖垮,而且你的客户端也可能因为请求堆积崩溃
优化后的批量插入方案(Node.js示例)
Elasticsearch提供了Bulk API,专门用来批量处理文档写入,这才是生成百万级数据的正确方式。下面是完整的代码示例,基于官方的@elastic/elasticsearch客户端:
const { Client } = require('@elastic/elasticsearch'); const client = new Client({ node: 'http://localhost:9200' }); // 提前创建测试索引,调整参数提升写入性能 async function createTestIndex() { try { await client.indices.create({ index: 'test', body: { settings: { number_of_shards: 3, // 根据你的节点数调整,建议每个分片不超过50GB number_of_replicas: 0, // 插入期间关闭副本,完成后再开启 refresh_interval: '-1' // 插入期间不自动刷新,大幅提升写入速度 }, mappings: { properties: { timestamp: { type: 'date' }, random_field: { type: 'keyword' }, numeric_field: { type: 'integer' } } } } }); console.log('测试索引创建成功'); } catch (err) { if (err.meta.statusCode === 400) { console.log('索引已存在,跳过创建'); } else { throw err; } } } // 生成批量插入的请求体 function generateBulkBatch(startId, batchSize) { const bulkBody = []; const now = Date.now(); for (let i = 0; i < batchSize; i++) { const docId = startId + i; // 生成模拟的测试数据,你可以根据需求调整字段 const timestamp = new Date(now - Math.floor(Math.random() * 30 * 24 * 60 * 60 * 1000)); // 近30天的随机时间 const randomValue = `test_value_${Math.floor(Math.random() * 1000000)}`; const numericValue = Math.floor(Math.random() * 1000); bulkBody.push({ index: { _index: 'test', _id: docId.toString() } }); bulkBody.push({ timestamp: timestamp.toISOString(), random_field: randomValue, numeric_field: numericValue }); } return bulkBody; } // 批量插入百万条数据 async function insertMillionRecords(totalCount = 1000000, batchSize = 1000) { await createTestIndex(); const totalBatches = Math.ceil(totalCount / batchSize); let successCount = 0; for (let batch = 0; batch < totalBatches; batch++) { const startId = batch * batchSize; const currentBatchSize = batch === totalBatches - 1 ? totalCount - startId : batchSize; const bulkBody = generateBulkBatch(startId, currentBatchSize); try { const response = await client.bulk({ body: bulkBody }); // 统计成功插入的数量 response.items.forEach(item => { if (item.index.status === 201) successCount++; }); console.log(`完成第${batch+1}/${totalBatches}批插入,累计成功${successCount}条`); } catch (err) { console.error(`第${batch+1}批插入失败:`, err.message); } } // 插入完成后,恢复索引的自动刷新和副本设置 await client.indices.putSettings({ index: 'test', body: { refresh_interval: '1s', number_of_replicas: 1 // 根据你的需求调整副本数 } }); console.log(`百万条数据插入完成,最终成功${successCount}条`); } // 执行插入 insertMillionRecords().catch(err => console.error('整体执行失败:', err));
关键优化点说明
- 批量提交:每1000条数据提交一次,平衡请求大小和性能,你可以根据ES的配置调整
batchSize(建议500-2000之间) - 索引参数优化:插入期间关闭副本、禁用自动刷新,能把写入性能提升数倍
- 模拟真实数据:生成包含日期、字符串、数值的多字段文档,更贴近真实业务场景,测试内存占用更准确
- 唯一ID:用自增ID确保每条记录都能被正确保存
内存测试的额外建议
- 插入完成后,执行一次手动刷新:
await client.indices.refresh({ index: 'test' }),让所有数据都被加载到内存 - 查看内存占用可以用ES的命令:
- 查看节点内存:
curl http://localhost:9200/_cat/nodes?v(关注heap.percent和heap.max字段) - 查看索引内存占用:
curl http://localhost:9200/_cat/indices?v(关注pri.store.size和store.size字段)
- 查看节点内存:
- 如果要测试极端内存情况,可以增加数组、嵌套对象等复杂字段,或者调整字段的分词方式(比如text字段会占用更多内存)
内容的提问来源于stack exchange,提问作者Palaniichuk Dmytro
相关产品推荐
相关产品推荐

