You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Windows 10环境下基于node-redis流式导入50万条CSV至Redis

Hey there! Let's get that 500k-record CSV imported into Redis efficiently using streaming—great call going with this approach to avoid memory overload. Below are two tailored solutions for Redis Set and Hash structures, building on your existing code.


Using Redis Set for Unique Entries

If you just need to store unique records (and don't need to access individual fields later), a Redis Set is perfect. We'll serialize each row to a JSON string (or use a unique field if you have one) and add it to the set:

const fs = require('fs');
const csv = require('csv-parser');
const redis = require('redis');
const client = redis.createClient();

client.on('connect', function() {
  console.log('Redis client connected');
});

client.on('error', function (err) {
  console.log('Something went wrong ' + err);
});

fs.createReadStream('test.csv')
  .pipe(csv())
  .on('data', async (row) => {
    // Serialize the entire row, or use a unique field like row.id instead
    const entry = JSON.stringify(row);
    try {
      await client.sAdd('csv_records_set', entry);
    } catch (err) {
      console.error('Failed to add entry to Set:', err);
    }
  })
  .on('end', async () => {
    console.log('CSV file successfully processed');
    // Don't forget to close the Redis connection!
    await client.quit();
  });

Note: If your CSV has a unique identifier column (like id), replace JSON.stringify(row) with row.id to save space and enforce uniqueness automatically.


Using Redis Hash for Structured Data

If you want to keep your data organized and access specific fields later, use Redis Hashes. Each row will be stored under a unique key (using a column like id), with CSV fields mapped to Hash key-value pairs:

const fs = require('fs');
const csv = require('csv-parser');
const redis = require('redis');
const client = redis.createClient();

client.on('connect', function() {
  console.log('Redis client connected');
});

client.on('error', function (err) {
  console.log('Something went wrong ' + err);
});

// Assume your CSV has a unique "id" column to use as the Redis key
fs.createReadStream('test.csv')
  .pipe(csv())
  .on('data', async (row) => {
    if (!row.id) {
      console.warn('Skipping row without an ID:', row);
      return;
    }
    const redisKey = `csv_record:${row.id}`;
    try {
      await client.hSet(redisKey, row);
    } catch (err) {
      console.error(`Failed to store record with ID ${row.id}:`, err);
    }
  })
  .on('end', async () => {
    console.log('CSV file successfully processed');
    await client.quit();
  });

Why this works: Later, you can fetch individual fields with client.hGet('csv_record:123', 'email') or the entire record with client.hGetAll('csv_record:123').


Speed Up Large Imports with Batching

Processing 500k rows one at a time can be slow. For better performance, batch your Redis operations:

const fs = require('fs');
const csv = require('csv-parser');
const redis = require('redis');
const client = redis.createClient();

const BATCH_SIZE = 1000;
let batch = [];

client.on('connect', function() {
  console.log('Redis client connected');
});

client.on('error', function (err) {
  console.log('Something went wrong ' + err);
});

fs.createReadStream('test.csv')
  .pipe(csv())
  .on('data', (row) => {
    batch.push(JSON.stringify(row));
    // Execute batch when we hit the size limit
    if (batch.length >= BATCH_SIZE) {
      client.sAdd('csv_records_set', batch)
        .catch(err => console.error('Batch import failed:', err))
        .finally(() => batch = []);
    }
  })
  .on('end', async () => {
    // Process any remaining records in the batch
    if (batch.length > 0) {
      await client.sAdd('csv_records_set', batch);
    }
    console.log('CSV file successfully processed');
    await client.quit();
  });

Adjust BATCH_SIZE based on your Redis server's capacity—1000-5000 is a safe starting point.


Quick Checks After Import

Verify everything worked with these Redis commands:

  • For Sets: SCARD csv_records_set (counts entries) or SMEMBERS csv_records_set (shows sample entries)
  • For Hashes: HGETALL csv_record:123 (fetches a specific record) or KEYS csv_record:* (lists all record keys)

内容的提问来源于stack exchange,提问作者Lokesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:10:55