PHP Curl大文件下载异常:CRONJOB无法保存15MB ZIP文件
问题描述
我编写了一个基于Curl的PHP函数,用于获取编码ZIP文件响应并保存至服务器,该脚本通过CRONJOB调用。下载1MB文件时一切正常;但下载15MB文件时,浏览器调用脚本可正常完成下载,通过CRONJOB执行时文件却完全未保存到服务器,且错误日志中无任何报错信息。浏览器下载该文件耗时约10秒,怀疑是CRONJOB存在超时相关问题?
原代码
function saveRemoteFile($url, $token, $type ) { // API Request // initialize cURL $ch = curl_init(); curl_setopt($ch, CURLOPT_URL,$url); curl_setopt($ch, CURLOPT_CUSTOMREQUEST, "GET"); curl_setopt($ch, CURLOPT_HTTPHEADER, array( 'accept: */*', 'Authorization: Bearer '.$token )); curl_setopt($ch, CURLOPT_TIMEOUT, 400); curl_setopt($ch, CURLOPT_CONNECTTIMEOUT ,0); curl_setopt($ch, CURLOPT_HEADER, true); curl_setopt($ch, CURLOPT_ENCODING, ''); curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1); // Execute cURL and store the response in a variable $file = curl_exec($ch); // Get the Header Size $header_size = curl_getinfo($ch, CURLINFO_HEADER_SIZE); // Get the Header from response $header = substr($file, 0, $header_size); // Get the Body from response $body = substr($file, $header_size); // Explode Header rows into an array $header_items = explode("\n", $header); // Close cURL handler curl_close($ch); // find the filname in the headers. $file_name = "BIHR.zip"; // Check header response, if HTTP response is not 200, then display the error. if(!preg_match('/200/', $header_items[0])){ echo '<pre>'.print_r($header_items[0], true).'</pre>'; exit(); } else { // Check header response, if HTTP response is 200, then proceed further. // Set the header for PHP to tell it, we would like to download a file header('Content-Description: File Transfer'); header('Content-Type: application/octet-stream'); header('Content-Transfer-Encoding: binary'); header('Expires: 0'); header('Cache-Control: must-revalidate'); header('Pragma: public'); header('Content-Disposition: attachment; filename='.$file_name); $destdir = '/******/file'; file_put_contents($destdir.'/'.date('y').date('m').date('d').'bihr_'.$type.'.zip',$body); exit; } }
问题分析与解决方案
1. PHP CLI模式超时限制
CRON执行PHP时使用的是CLI版配置,和网页环境的php.ini参数可能不同。默认CLI的max_execution_time可能较短(比如30秒),即便Curl设置了400秒超时,PHP本身的执行超时会直接终止脚本。
- 解决:在脚本开头添加
set_time_limit(0);,或者修改CLI版php.ini的max_execution_time为0(无限制)。
2. 内存占用过高
原脚本把整个15MB文件加载到内存中($file变量),再拆分header和body,CRON环境的内存限制可能更低,导致内存不足直接终止脚本且无日志输出。
- 解决:改用流式下载,边下载边写入文件,避免占用大量内存:
function saveRemoteFile($url, $token, $type ) { $destDir = '/******/file'; $destPath = $destDir . '/' . date('ymd') . 'bihr_' . $type . '.zip'; // 检查目标目录是否可写 if (!is_dir($destDir) || !is_writable($destDir)) { error_log("目标目录不可写或不存在: $destDir"); exit; } $fp = fopen($destPath, 'w'); if (!$fp) { error_log("无法打开目标文件: $destPath"); exit; } $ch = curl_init(); curl_setopt($ch, CURLOPT_URL, $url); curl_setopt($ch, CURLOPT_CUSTOMREQUEST, "GET"); curl_setopt($ch, CURLOPT_HTTPHEADER, array( 'accept: */*', 'Authorization: Bearer ' . $token )); curl_setopt($ch, CURLOPT_TIMEOUT, 400); curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 0); curl_setopt($ch, CURLOPT_HEADER, false); // 单独获取响应码,不把header写入文件 curl_setopt($ch, CURLOPT_ENCODING, ''); curl_setopt($ch, CURLOPT_FILE, $fp); // 直接将响应写入本地文件 curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true); // 自动跟随重定向 curl_exec($ch); // 检查Curl执行错误 if (curl_errno($ch)) { $errorMsg = curl_error($ch); error_log("Curl执行错误: $errorMsg"); fclose($fp); unlink($destPath); // 删除未完成的文件 exit; } // 检查HTTP响应码 $httpCode = curl_getinfo($ch, CURLINFO_HTTP_CODE); if ($httpCode !== 200) { error_log("HTTP响应错误,状态码: $httpCode"); fclose($fp); unlink($destPath); exit; } curl_close($ch); fclose($fp); }
3. 目录权限问题
CRON执行的用户(如root、cron专用用户)可能和web服务器用户(如www-data)不同,导致没有目标目录的写入权限。
- 解决:检查
/******/file目录的权限,确保CRON执行用户拥有读写权限;或者在脚本中添加is_writable判断,失败时记录日志。
4. 错误日志配置问题
CRON的错误日志默认可能不输出到你查看的网页错误日志中,导致误以为无报错。
- 解决:修改CRON命令,指定日志输出路径,比如:
* * * * * php /path/to/your/script.php >> /var/log/cron_script.log 2>&1
这样所有输出和错误都会写入指定日志,方便排查问题。
内容的提问来源于stack exchange,提问作者Steve Pasco
相关产品推荐
相关产品推荐

