PHP脚本收邮件存MySQL:发件人地址与正文处理问题
嘿,这两个问题我都踩过坑,给你分享下靠谱的解决办法!
问题1:提取发件人邮箱地址而非显示名称
你现在的代码是直接把From:字段的全部内容取出来,比如遇到From: "张三" <zhangsan@example.com>这种格式,就会拿到整个字符串,而不是里面的邮箱。我们需要用更精准的正则来匹配尖括号里的邮箱地址,同时兼容没有名称的情况(比如From: lisi@example.com)。
修改后的代码:
// handle email $lines = explode("\n", $email); // empty vars $from = ""; $subject = ""; $headers = ""; $message = ""; $splittingheaders = true; for ($i=0; $i < count($lines); $i++) { if ($splittingheaders) { // this is a header $headers .= $lines[$i]."\n"; // look out for special headers if (preg_match("/^Subject: (.*)/", $lines[$i], $matches)) { $subject = $matches[1]; } // 处理发件人地址:优先匹配带尖括号的邮箱,再匹配直接的邮箱 if (preg_match("/^From:.*<([a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,})>/", $lines[$i], $matches)) { $from = $matches[1]; } elseif (preg_match("/^From: ([a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,})/", $lines[$i], $matches)) { $from = $matches[1]; } else { // 兜底:如果匹配不到邮箱,就用原内容(比如发件人只有名称的极端情况) $from = trim($matches[1] ?? $lines[$i]); } if (preg_match("/^To: (.*)/", $lines[$i], $matches)) { $to = $matches[1]; } } else { // content/main message body information $message .= $lines[$i]."\n"; } if (trim($lines[$i])=="") { // empty line, header section has ended $splittingheaders = false; } }
这个正则会准确提取出邮箱地址,不管发件人字段有没有显示名称。
问题2:清理邮件正文的MIME冗余内容
你原来的方法是用空行分割邮件头和正文,但现代邮件大多是MIME多格式的(同时包含纯文本和HTML),所以正文里会带MIME分隔符、Content-Type这些头信息。手动用正则替换是不靠谱的,最好用专门的邮件解析工具来处理。
推荐方案:使用mailparse扩展(最可靠)
PHP的mailparse扩展是专门用来解析邮件的,能轻松处理各种MIME结构。如果你的服务器没装这个扩展,可以联系主机商安装,或者通过PECL安装(pecl install mailparse)。
整合到你的代码里的示例:
// 先处理完邮件头(用上面修改后的代码拿到$headers和$message的原始内容) // 解析邮件内容,提取纯文本正文 $resource = mailparse_msg_create(); mailparse_msg_parse($resource, $email); // $email是完整的原始邮件内容 $structure = mailparse_msg_get_structure($resource); $clean_message = ''; foreach ($structure as $part_id) { $part = mailparse_msg_get_part($resource, $part_id); $part_data = mailparse_msg_get_part_data($part); // 只提取text/plain类型的正文(如果要HTML正文就找text/html) if (strpos($part_data['content-type'], 'text/plain') !== false) { // 提取该部分内容到临时流 mailparse_msg_extract_part_file($part, 'php://temp'); $temp_stream = fopen('php://temp', 'r'); $clean_message = stream_get_contents($temp_stream); fclose($temp_stream); // 处理编码(比如quoted-printable) if (isset($part_data['transfer-encoding']) && strtolower($part_data['transfer-encoding']) === 'quoted-printable') { $clean_message = quoted_printable_decode($clean_message); } break; // 找到第一个纯文本部分就停止,避免重复 } } mailparse_msg_free($resource); // 现在$clean_message就是干净的纯文本正文了 echo $clean_message;
备选方案:手动解析MIME(无扩展时用)
如果没法装mailparse,可以手动处理简单的多部分邮件,但兼容性差一些:
// 先从邮件头里找到MIME边界 preg_match("/^Content-Type: multipart\/.*boundary=(.*)/mi", $headers, $boundary_matches); $clean_message = $message; if (!empty($boundary_matches[1])) { $boundary = trim($boundary_matches[1], '"'); $parts = explode("--".$boundary, $clean_message); foreach ($parts as $part) { // 找text/plain的部分 if (strpos($part, 'Content-Type: text/plain') !== false) { // 用空行分割该部分的头和内容 $part_content = explode("\n\n", $part, 2); if (isset($part_content[1])) { $clean_message = trim($part_content[1]); // 处理编码 if (strpos($part, 'Content-Transfer-Encoding: quoted-printable') !== false) { $clean_message = quoted_printable_decode($clean_message); } break; } } } } else { // 非多部分邮件,直接处理编码 if (strpos($headers, 'Content-Transfer-Encoding: quoted-printable') !== false) { $clean_message = quoted_printable_decode($clean_message); } } // $clean_message就是处理后的正文
这个方法能应付大多数普通邮件,但遇到嵌套多部分或者复杂编码的邮件可能会失效,所以优先推荐mailparse。
内容的提问来源于stack exchange,提问作者Charlie
相关产品推荐
相关产品推荐

