You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Chrome Extension程序化读取Gmail的PDF附件内容?

Hey there! I’ve built Chrome Extensions that interact with Gmail before, so let me break down exactly how you can programmatically read the content of PDF attachments from Gmail. Here's a step-by-step guide:

How to Read Gmail PDF Attachments in a Chrome Extension

1. Set Up Required Permissions in manifest.json

First, you need to declare the right permissions in your extension's manifest to access Gmail and handle scripts/data. For Manifest V3, your manifest.json should look something like this:

{
  "manifest_version": 3,
  "name": "Gmail PDF Parser",
  "version": "1.0",
  "permissions": ["gmail", "scripting", "identity"],
  "host_permissions": ["https://mail.google.com/*"],
  "oauth2": {
    "client_id": "YOUR_GOOGLE_CLOUD_CLIENT_ID.apps.googleusercontent.com",
    "scopes": ["https://www.googleapis.com/auth/gmail.readonly"]
  }
}

Note: You'll need to create a project in Google Cloud Console, enable the Gmail API, and generate a client ID for the OAuth setup.

2. Fetch the Target Email & Identify PDF Attachments

Use the Gmail API to pull email details, then filter out the PDF attachments. Here's a code snippet to get started:

// Get OAuth token to authenticate API calls
chrome.identity.getAuthToken({ interactive: true }, (token) => {
  if (chrome.runtime.lastError) {
    console.error(chrome.runtime.lastError.message);
    return;
  }

  // Replace {MESSAGE_ID} with the ID of the email you want to process
  fetch(`https://www.googleapis.com/gmail/v1/users/me/messages/{MESSAGE_ID}?format=full`, {
    headers: { Authorization: `Bearer ${token}` }
  })
  .then(res => res.json())
  .then(message => {
    // Filter for PDF attachments
    const pdfAttachments = message.payload.parts.filter(part => 
      part.filename && part.filename.toLowerCase().endsWith('.pdf')
    );

    // Process each PDF attachment
    pdfAttachments.forEach(attachment => {
      const attachmentId = attachment.body.attachmentId;
      fetchPDFAttachment(token, message.id, attachmentId);
    });
  });
});

3. Retrieve the PDF's Binary Content

Once you have the attachmentId, call the Gmail API to get the attachment data (which comes in base64 encoding). Convert it to a Blob for parsing:

function fetchPDFAttachment(token, messageId, attachmentId) {
  fetch(`https://www.googleapis.com/gmail/v1/users/me/messages/${messageId}/attachments/${attachmentId}`, {
    headers: { Authorization: `Bearer ${token}` }
  })
  .then(res => res.json())
  .then(attachmentData => {
    // Decode base64 data (Gmail uses URL-safe base64, so replace chars first)
    const decodedData = atob(attachmentData.data.replace(/-/g, '+').replace(/_/g, '/'));
    const byteArray = new Uint8Array(decodedData.length);
    
    for (let i = 0; i < decodedData.length; i++) {
      byteArray[i] = decodedData.charCodeAt(i);
    }

    // Create a Blob of the PDF
    const pdfBlob = new Blob([byteArray], { type: 'application/pdf' });
    extractTextFromPDF(pdfBlob);
  });
}

4. Parse the PDF Content with PDF.js

To read the text from the PDF Blob, use Mozilla's PDF.js library (it's the most reliable tool for this). First, include pdf.js and pdf.worker.js in your extension's files, then use this code:

function extractTextFromPDF(pdfBlob) {
  // Configure PDF.js worker
  pdfjsLib.GlobalWorkerOptions.workerSrc = chrome.runtime.getURL('pdf.worker.js');

  const reader = new FileReader();
  reader.onload = () => {
    const arrayBuffer = reader.result;
    pdfjsLib.getDocument(arrayBuffer).promise
    .then(pdf => {
      let fullText = '';
      // Loop through each page to extract text
      const pagePromises = [];
      for (let pageNum = 1; pageNum <= pdf.numPages; pageNum++) {
        pagePromises.push(
          pdf.getPage(pageNum).then(page => page.getTextContent())
        );
      }

      // Wait for all pages to process
      Promise.all(pagePromises).then(pageContents => {
        pageContents.forEach(content => {
          content.items.forEach(item => {
            fullText += `${item.str} `;
          });
        });
        // Now you have the full text of the PDF!
        console.log('Extracted PDF Text:', fullText);
        // Do whatever you need with the text here (save, display, etc.)
      });
    })
    .catch(err => console.error('Error parsing PDF:', err));
  };

  reader.readAsArrayBuffer(pdfBlob);
}

Key Notes to Keep in Mind

  • OAuth Consent: Make sure your Google Cloud project's OAuth consent screen is properly configured (for public extensions, you'll need to get it verified by Google).
  • Alternative: DOM Injection: If you don't want to use the Gmail API, you could inject a script into the Gmail web page to find PDF attachment links, download them, and parse. But this is fragile—Gmail's DOM structure changes often, so the API is the better long-term solution.
  • Permissions: Always request the minimal permissions needed (we used gmail.readonly here since we don't need to modify emails).

内容的提问来源于stack exchange,提问作者Sabarish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:35:30