基于googleapiclient的Python邮件解析与线程回复问题
Hey there, let's break down how to solve your two key challenges: reliably extracting the user's actual reply content (excluding old email quotes) and sending a properly threaded response that follows RFC 2822 standards.
1. Extracting the User's Reply Content
The core issue here is separating the user's new message from the quoted old conversation. Most email clients (including Gmail) prefix quoted content with a line starting with On [date], [sender] wrote: followed by the quoted lines (marked with >). We can use regex to target this separator reliably.
Solution Code
import re from bs4 import BeautifulSoup # Only needed if handling HTML emails import base64 def extract_user_reply(decoded_content, content_type="text/plain"): # Convert bytes to string if needed if isinstance(decoded_content, bytes): content_str = decoded_content.decode('utf-8') else: content_str = decoded_content # Handle HTML content by converting to plain text first if content_type == "text/html": soup = BeautifulSoup(content_str, 'html.parser') content_str = soup.get_text() # Regex pattern to match the start of quoted old emails # Matches lines starting with "On " followed by date/sender and "wrote:", plus subsequent empty lines quote_separator_pattern = r'On\s.*wrote:\r?\n\r?\n' match = re.search(quote_separator_pattern, content_str) if match: # Grab everything before the quote separator reply_content = content_str[:match.start()] else: # No quote found - use the entire content (likely a reply without quotes) reply_content = content_str # Clean up extra newlines and whitespace return reply_content.strip() # Example usage with a thread message def get_latest_thread_content(service, thread_id): thread = service.users.threads().get(userId='me', id=thread_id).execute() latest_message = thread['messages'][-1] payload = latest_message['payload'] # Find the plain text or HTML part of the email content = "" content_type = "text/plain" if 'parts' in payload: for part in payload['parts']: if part['mimeType'] in ['text/plain', 'text/html']: content_type = part['mimeType'] content = base64.urlsafe_b64decode(part['body']['data']) break else: content = base64.urlsafe_b64decode(payload['body']['data']) return extract_user_reply(content, content_type)
Key Notes
- The regex handles both
\r\n(Windows) and\n(Unix) line endings. - If the user's reply doesn't include any quoted content, the function returns the entire message (after cleaning whitespace).
- For HTML emails, we use
BeautifulSoupto convert HTML to plain text first—install it withpip install beautifulsoup4if needed.
2. Sending a Threaded Reply (RFC 2822 Compliant)
To ensure your reply lands in the original thread, you need to set two critical email headers: In-Reply-To and References, plus adjust the subject line (add Re: if missing). You also need to include the original thread ID in your send request for extra reliability.
Solution Code
from email.mime.text import MIMEText import base64 def create_thread_reply(sender, to, original_subject, original_message_id, thread_id, message_text): # Add "Re:" prefix to subject if it doesn't already exist if not original_subject.startswith('Re: '): subject = f'Re: {original_subject}' else: subject = original_subject # Build the MIME message message = MIMEText(message_text) message['to'] = to message['from'] = sender message['subject'] = subject # RFC 2822 required headers for threading message['In-Reply-To'] = original_message_id message['References'] = original_message_id # Encode the message and include the thread ID raw_message = base64.urlsafe_b64encode(message.as_bytes()).decode() return {'raw': raw_message, 'threadId': thread_id} # Example workflow to send the reply def send_threaded_reply(service, thread_id, user_reply): # Get original thread details thread = service.users.threads().get(userId='me', id=thread_id).execute() latest_message = thread['messages'][-1] # Extract original message ID (from the user's reply email) original_message_id = None for header in latest_message['payload']['headers']: if header['name'] == 'Message-ID': original_message_id = header['value'] break # Extract original subject (from the first email in the thread) original_subject = None for header in thread['messages'][0]['payload']['headers']: if header['name'] == 'Subject': original_subject = header['value'] break # Configure your reply details sender = 'your-email@gmail.com' to = 'recipient-email@gmail.com' # Or extract from the thread's headers reply_text = f"Thanks for your reply! You said: {user_reply}" # Create and send the reply reply_message = create_thread_reply(sender, to, original_subject, original_message_id, thread_id, reply_text) service.users.messages().send(userId='me', body=reply_message).execute()
Key Notes
In-Reply-Toshould contain the Message-ID of the email you're directly replying to (in this case, the user's reply message).Referencesshould include the Message-ID of the original thread's root email (or the full chain of Message-IDs—using just the root works for Gmail).- Including
threadIdin the send request ensures Gmail explicitly ties your reply to the thread, even if headers are slightly off.
Full Workflow Example
# Initialize your Gmail API service first (follow Google's setup guide) # service = build('gmail', 'v1', credentials=credentials) # 1. Extract user's reply thread_id = 'your-thread-id-here' user_reply = get_latest_thread_content(service, thread_id) print(f"User's reply: {user_reply}") # 2. Send threaded reply send_threaded_reply(service, thread_id, user_reply)
内容的提问来源于stack exchange,提问作者JolonB

