使用Zend Mail通过IMAP连接Gmail获取邮件原始内容的问题
Great question! Parsing raw MIME email content manually (splitting by boundaries, decoding encoding schemes) is error-prone—especially with nested MIME structures or edge cases like quoted-printable encoding. Since you're already using Zend Mail, let's leverage its built-in MIME handling tools to cleanly extract and process the plain text, HTML, links, and images from your raw email content.
Step-by-Step Solution
Here's a complete code example that parses the raw content, extracts each MIME part, and processes key elements like text, links, and images:
use Zend\Mime\Message as MimeMessage; use Zend\Mime\Part as MimePart; // Get your raw email content (from your existing code) $rawContent = $mail->getRawContent($id); // Parse the raw content into a Zend MimeMessage object $mimeMessage = MimeMessage::createFromString($rawContent); // Iterate through each MIME part in the message foreach ($mimeMessage->getParts() as $part) { /** @var MimePart $part */ $contentType = $part->getHeaders()->get('Content-Type')->getFieldValue(); $decodedContent = $part->getContent(); // Automatically handles quoted-printable/base64 encoding // Process plain text part if (strpos($contentType, 'text/plain') !== false) { echo "### Plain Text Content\n"; echo $decodedContent . "\n\n"; // Extract links from plain text preg_match_all('/https?:\/\/[^\s]+/', $decodedContent, $plainLinks); if (!empty($plainLinks[0])) { echo "#### Plain Text Links\n"; foreach ($plainLinks[0] as $link) { // Fix quoted-printable encoding artifacts (=3D is the encoded '=') $cleanLink = str_replace('=3D', '=', $link); echo "- " . $cleanLink . "\n"; } echo "\n"; } // Extract image alt text from plain text markers preg_match_all('/\[图片: ([^\]]+)\]/', $decodedContent, $plainImages); if (!empty($plainImages[1])) { echo "#### Plain Text Image Alt Text\n"; foreach ($plainImages[1] as $alt) { echo "- " . $alt . "\n"; } echo "\n"; } } // Process HTML part elseif (strpos($contentType, 'text/html') !== false) { echo "### HTML Content\n"; echo $decodedContent . "\n\n"; // Parse HTML to extract links and images (use DOMDocument for reliability) $dom = new DOMDocument(); // Suppress HTML parsing warnings for malformed email HTML libxml_use_internal_errors(true); $dom->loadHTML($decodedContent); libxml_clear_errors(); // Extract HTML links $links = $dom->getElementsByTagName('a'); if ($links->length > 0) { echo "#### HTML Links\n"; foreach ($links as $a) { $href = str_replace('=3D', '=', $a->getAttribute('href')); $linkText = trim($a->textContent); echo "- Text: *" . $linkText . "* | URL: `" . $href . "`\n"; } echo "\n"; } // Extract HTML images $images = $dom->getElementsByTagName('img'); if ($images->length > 0) { echo "#### HTML Images\n"; foreach ($images as $img) { $src = str_replace('=3D', '=', $img->getAttribute('src')); $alt = $img->getAttribute('alt'); echo "- Alt Text: *" . $alt . "* | Image URL: `" . $src . "`\n"; } echo "\n"; } } }
Key Notes
- Avoid manual MIME boundary splitting: Zend's
MimeMessagehandles nested MIME structures (like multipart/alternative or multipart/mixed) automatically, so you don't have to worry about edge cases. - Automatic encoding handling: The
getContent()method onMimePartdecodes quoted-printable, base64, and other common transfer encodings out of the box. - Fix quoted-printable artifacts: You'll notice links in the raw content have
=3Dinstead of=—this is a quoted-printable encoding quirk, so we replace it to get valid URLs. - Reliable HTML parsing: Using
DOMDocumentinstead of regular expressions ensures you correctly extract links and images even with malformed email HTML (which is common).
Alternative: Using PHP's Built-in IMAP Functions
If you prefer to use PHP's native IMAP functions instead of Zend components, you can use imap_fetchstructure() to get the email structure and imap_body() to retrieve individual parts. However, this requires more manual handling of encoding and structure traversal compared to Zend's solution.
内容的提问来源于stack exchange,提问作者GrahamMorbyDev

