PHP通过Shell与本地Web服务器运行时行为差异排查求助
问题排查:CalDAV日历ICS下载终端运行异常
我开发了一个项目,用于下载指定CalDAV日历链接的内容并合并ICS文件。但使用file_put_contents或file_get_contents时出现异常:通过本地Web服务器运行能获取正确的ICS内容,通过Shell终端运行却下载了网页内容而非预期的ICS文件。
获取到的网页代码片段
<!DOCTYPE html> <html> <head> <title>public-calendars - sabre/dav </title> <link rel="shortcut icon" href="/remote.php/dav/?sabreAction=asset&assetName=favicon.ico" type="image/vnd.microsoft.icon" /> <link rel="stylesheet" href="/remote.php/dav/?sabreAction=asset&assetName=sabredav.css" type="text/css" /> <link rel="stylesheet" href="/remote.php/dav/?sabreAction=asset&assetName=openiconic%2Fopen-iconic.css" type="text/css" /> </head> <body> <header> <div class="logo"> <a href="/remote.php/dav/"><img src="/remote.php/dav/?sabreAction=asset&assetName=sabredav.png" alt="sabre/dav" /> public-calendars/ysT7xJQmSxx2RTKZ</a> </div> </header> <nav><a href="/remote.php/dav/public-calendars" class="btn">⇤ Go to parent</a> <a href="?sabreAction=plugins" class="btn"><span class="oi" data-glyph="puzzle-piece"></span> Plugins</a></nav><section><h1>Nodes</h1> <table class="nodeTable"><tr><td class="nameColumn"><a href="/remote.php/dav/public-calendars/ysT7xJQmSxx2RTKZ/0A4D04A5-90FF-477B-A8DC-C703E1FACA1F.ics"><span class="oi" data-glyph="file"></span> 0A4D04A5-90FF-477B-A8DC-C703E1FACA1F.ics</a></td><td class="typeColumn">File</td><td>447 bytes</td><td>May 1, 2022, 5:15 pm</td><td></td><td><a href="/remote.php/dav/public-calendars/ysT7xJQmSxx2RTKZ/0A4D04A5-90FF-477B-A8DC-C703E1FACA1F.ics?sabreAction=info"><span class="oi" data-glyph="info"></span></a></td></tr><tr><td class="nameColumn"><a /// 内容过长已裁剪
原PHP代码
<?php function processFile($i, $size, $file) { file_put_contents('test.ics', file_get_contents($file, FALSE, NULL)); $lines = file($file); $beginEvent = array_search("BEGIN:VEVENT", $lines); $reverselines = array_reverse($lines, true); $endEvent = array_search('END:VEVENT', $reverselines); if ($size == 1) { // do nothing } elseif ($i == 0) { //First calendar file entry: Takes in also first part of meta data $lines = array_slice($lines, 0, $endEvent + 1); } elseif ($i == $size - 1) { //Last calendar file entry: Take last part of meta data $lines = array_slice($lines, $beginEvent); } else { // Slice only Event information, beginnen and ending with "VEVENT" headers $lines = array_slice($lines, $beginEvent, $endEvent - $beginEvent + 1); } return $lines; } $links = file("links.txt"); $size = count($links) - 1; $result = array(); $tempresult = array(); for ($i = 0; $i < $size; $i++) { $tempresult = processFile($i, $size, $links[$i]); //print_r($tempresult); $result = array_merge($result, $tempresult); } $file = fopen("output.ics", "w"); foreach ($result as $value) { fwrite($file, $value . " "); } fclose($file); ?>
问题原因
终端运行的PHP和Web服务器环境的核心差异在于请求头中的User-Agent:
- Web环境下,请求由浏览器发起,User-Agent是浏览器标识,CalDAV服务器(此处为sabre/dav)会返回ICS文件内容;
- 终端运行PHP时,
file_get_contents默认使用的User-Agent是PHP/[版本号],服务器识别为非浏览器请求,返回网页版的目录列表,也就是你看到的HTML内容。
另外,file("links.txt")读取的每一行链接会包含换行符(\n或\r\n),直接传递给file_get_contents会导致请求URL无效,也是可能的诱因。
修复方案
1. 给请求添加浏览器User-Agent
修改file_get_contents调用,添加上下文参数,模拟浏览器请求:
// 构造请求上下文,设置浏览器User-Agent $context = stream_context_create([ 'http' => [ 'header' => "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36\r\n" ] ]); // 使用上下文发起请求,同时清理链接中的换行符 $icsContent = file_get_contents(trim($file), FALSE, $context); file_put_contents('test.ics', $icsContent); $lines = explode("\n", $icsContent); // 直接从内容拆分行,避免重复读取文件
2. 清理链接中的换行符
读取links.txt时,对每个链接做trim()处理,去掉多余的换行和空格:
$links = array_map('trim', file("links.txt")); $size = count($links); // 清理后无需再减1,原代码减1是因为读取的空行被算入,现在已过滤
完整修改后的代码
<?php function processFile($i, $size, $file) { // 构造浏览器请求上下文 $context = stream_context_create([ 'http' => [ 'header' => "User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36\r\n" ] ]); $icsContent = file_get_contents(trim($file), FALSE, $context); file_put_contents('test.ics', $icsContent); $lines = explode("\n", $icsContent); $beginEvent = array_search("BEGIN:VEVENT", $lines); $reverselines = array_reverse($lines, true); $endEvent = array_search('END:VEVENT', $reverselines); if ($size == 1) { // 单个文件无需处理 } elseif ($i == 0) { // 第一个文件:保留到第一个END:VEVENT为止的元数据和事件 $lines = array_slice($lines, 0, $endEvent + 1); } elseif ($i == $size - 1) { // 最后一个文件:从最后一个BEGIN:VEVENT开始保留事件和元数据 $lines = array_slice($lines, $beginEvent); } else { // 中间文件:只保留BEGIN:VEVENT到END:VEVENT的事件内容 $lines = array_slice($lines, $beginEvent, $endEvent - $beginEvent + 1); } return $lines; } $links = array_map('trim', file("links.txt")); $size = count($links); $result = array(); for ($i = 0; $i < $size; $i++) { $tempresult = processFile($i, $size, $links[$i]); $result = array_merge($result, $tempresult); } $file = fopen("output.ics", "w"); foreach ($result as $value) { fwrite($file, $value . "\n"); } fclose($file); ?>
内容的提问来源于stack exchange,提问作者Finn Bonnen
相关产品推荐
相关产品推荐

