You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

因内存限制无法加载50万条数据到数组,求优化方案

解决50万条数据解析内存溢出问题

我正在将约50万条条目解析到数组$properties中,代码如下:

$properties = array(); 
$handle = fopen($file_path, "r"); 
if ($handle) { 
    while (($str = fgets($handle)) !== false) { 
        if (strlen($str) && $str[0] == '#') { 
            $pdate = substr($str, 1); 
            $date = rtrim($pdate); 
            $formatted = DateTime::createFromFormat('* M d H:i:s T Y', $date); 
        } 
        $str = rtrim ($str, "\n"); 
        $exp = explode ('=', $str); 
        if (count($exp) == 2){ 
            $exp2 = explode('.', $exp[0]); 
            if (count($exp2) == 2) { 
                if ($exp2[1] == "dateTime") { 
                    $s = str_replace("\\", "", $exp[1]); 
                    $d = strtotime($s); 
                    $dateTime = date('Y-m-d H:i:s', $d); 
                    $properties[$exp2[0]][$exp2[1]] = $dateTime; 
                } else { 
                    $properties[$exp2[0]][$exp2[1]] = $exp[1]; 
                } 
            } else { 
                $properties[$exp[0]] = $exp[1]; 
            } 
        } 
    } 
    fclose($handle); 
} else { 
    echo "error"; 
}

目前代码可正常运行,但因数组过大无法处理,我尝试使用以下代码拆分数组:

$properties_chunk = array_chunk($properties, 10000, true); 

但系统崩溃,无法生成$properties_chunk数组。期望最终数组结构如下:

array(4) { 
    [0]=> array(10000) { 
        ["12345"]=> array(5) { 
            ["dateTime"]=> string(19) "2016-10-12 19:46:25" 
            ["fileName"]=> string(46) "monkey.jpg" 
            ["path"]=> string(149) "Volumes/animals/monkey.jpg" 
            ["size"]=> string(7) "2650752" 
        } 
        ["678790"]=> array(5) { 
            ["dateTime"]=> string(19) "2016-10-12 14:39:43" 
            ["fileName"]=> string(45) "elephant.jpg" 
            ["path"]=> string(171) "Volumes/animals/elephant.jpg" 
            ["size"]=> string(7) "2306688" 
        } 
        ... and so on. 
    }
}

请问该如何解决此问题?


问题根源

你现在的核心问题是一次性把50万条数据全部加载到内存中,这直接导致了内存溢出——哪怕后续用array_chunk拆分也没用,因为$properties已经占满了所有可用内存,系统根本没有空间再生成拆分后的数组。

解决方案:边解析边分批处理

核心思路是不要把所有数据都存在一个大数组里,而是每解析够10000条就立即处理(比如写入数据库、保存到文件),然后清空当前的小数组,继续解析下一批。这样内存里始终只保留一批数据,从根源上避免内存溢出。

修改后的代码示例

$batchSize = 10000;
$currentBatch = [];
$handle = fopen($file_path, "r");

if ($handle) {
    while (($str = fgets($handle)) !== false) {
        // 处理日期行(原代码中$formatted未使用,可根据实际需求调整逻辑)
        if (strlen($str) && $str[0] == '#') {
            $pdate = substr($str, 1);
            $date = rtrim($pdate);
            $formatted = DateTime::createFromFormat('* M d H:i:s T Y', $date);
            continue;
        }

        $str = rtrim($str, "\n");
        $exp = explode('=', $str);
        if (count($exp) != 2) {
            continue;
        }

        $exp2 = explode('.', $exp[0]);
        if (count($exp2) == 2) {
            $id = $exp2[0];
            $key = $exp2[1];
            
            // 初始化当前条目数组(如果不存在)
            if (!isset($currentBatch[$id])) {
                $currentBatch[$id] = [];
            }
            
            if ($key == "dateTime") {
                $s = str_replace("\\", "", $exp[1]);
                $d = strtotime($s);
                $dateTime = date('Y-m-d H:i:s', $d);
                $currentBatch[$id][$key] = $dateTime;
            } else {
                $currentBatch[$id][$key] = $exp[1];
            }
        } else {
            // 处理非二级结构的键值对
            $currentBatch[$exp[0]] = $exp[1];
        }

        // 检查当前批次是否达到设定大小,达到则处理并清空
        if (count($currentBatch) >= $batchSize) {
            processBatch($currentBatch);
            $currentBatch = [];
        }
    }

    // 处理最后一批不足10000条的数据
    if (!empty($currentBatch)) {
        processBatch($currentBatch);
    }

    fclose($handle);
} else {
    echo "error opening file";
}

// 自定义批次处理函数,根据你的需求实现(比如写入数据库、保存文件)
function processBatch($batch) {
    // 示例:打印批次大小,实际可替换为业务逻辑
    echo "Processing batch of " . count($batch) . " items\n";
    // 比如写入数据库:saveToDatabase($batch);
    // 或者保存为文件:file_put_contents('batch_' . uniqid() . '.php', '<?php return ' . var_export($batch, true) . ';');
}

关键优化点

  • 分批处理:内存始终只保留一个10000条的小批次,彻底解决内存溢出问题
  • 避免大数组:完全不需要生成$properties这个50万条的巨型数组,从根源上减少内存占用
  • 灵活扩展:processBatch函数可根据需求自定义——无论是写入数据库、生成多个小文件,都能实现你期望的分块数组结构

额外建议

  1. 如果临时需要调整内存限制,可以通过ini_set('memory_limit', '256M');增大,但这只是治标不治本,分批处理才是长久之计
  2. 原代码中$formatted变量未被使用,建议确认其用途,若不需要可删除以减少不必要的计算

内容的提问来源于stack exchange,提问作者peace_love

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:21:59