Python邮件脚本故障:Email列表所有项为同一值
问题
我写了一个脚本遍历邮件收件箱,找出含指定关键词的邮件,把发件人姓名、邮箱、公司信息分别存入列表,最后导入Pandas DataFrame并导出为CSV。现在姓名和公司列都正常,但邮箱地址列所有项都是同一个值,无法正确存储不同邮箱。
预期邮箱列表:
123@ABC.com 234@boo.com info@example.com
实际得到的邮箱列表:
123@ABC.com 123@ABC.com 123@ABC.com
完整代码:
import email import imaplib import re import pandas as pd EMAIL = "email" PASSWORD = 'password' SERVER = "imap.example.com" sender_list = [] sender_names = [] Sender_company = [] Sender_phone = {} x = 0 y = 1 # connect to the server and go to its inbox mail = imaplib.IMAP4_SSL(SERVER) sock=mail.socket() timeout = 60 * 20 # 5 minutes sock.settimeout(timeout) mail.login(EMAIL, PASSWORD) # we choose the inbox but you can select others mail.select('inbox') # we'll search using the ALL criteria to retrieve # every message inside the inbox # it will return with its status and a list of ids status, data = mail.search(None, 'ALL') # the list returned is a list of bytes separated # by white spaces on this format: [b'1 2 3', b'4 5 6'] # so, to separate it first we create an empty list mail_ids = [] # then we go through the list splitting its blocks # of bytes and appending to the mail_ids list for block in data: # the split function called without parameter # transforms the text or bytes into a list using # as separator the white spaces: # b'1 2 3'.split() => [b'1', b'2', b'3'] mail_ids += block.split() # now for every id we'll fetch the email # to extract its content for i in mail_ids: # the fetch function fetch the email given its id # and format that you want the message to be status, data = mail.fetch(i, '(RFC822)') # the content data at the '(RFC822)' format comes on # a list with a tuple with header, content, and the closing # byte b')' for response_part in data: # so if its a tuple... if isinstance(response_part, tuple): # we go for the content at its second element # skipping the header at the first and the closing # at the third message = email.message_from_bytes(response_part[1]) # with the content we can extract the info about # who sent the message and its subject mail_from = message['from'] mail_subject = message['subject'] sender_formatted = mail_from.replace(">","").split("<") list_length = len(sender_formatted) if list_length != 2: continue sender_email = sender_formatted[1] sender_name = sender_formatted[0] if sender_list.count(sender_email) > 0 : continue # then for the text we have a little more work to do # because it can be in plain text or multipart # if its not plain text we need to separate the message # from its annexes to get the text if message.is_multipart(): mail_content = '' # on multipart we have the text message and # another things like annex, and html version # of the message, in that case we loop through # the email payload for part in message.get_payload(): # if the content type is text/plain # we extract it if part.get_content_type() == 'text/plain': mail_content += part.get_payload() else: # if the message isn't multipart, just extract it mail_content = message.get_payload() words=["MF116","MF115","install","removal","U6","U16","U25","U40","U65","U100","meter","installed"] content= mail_content # print(content) for word in words: result = re.search(word,content) if result : # print (f"search success: {word}") sender_list.append(sender_email) sender_names.append(sender_name) company = sender_email.split("@") company = company[1].split(".") company = company[0] Sender_company.append(company) x += 1 if x == 100 : df = pd.DataFrame() df['Contact Name'] = sender_names df['Email Address'] = sender_email df['Company'] = Sender_company file_name = (f'contacts{y}.csv') df.to_csv(file_name,encoding='utf-8') y+1 # print(sender_email,company, sender_name) break else : # print(f"Search failed: {word}") continue # and then let's show its result # print(f'From: {mail_from}') # print(f'Subject: {mail_subject}') # print(f'Content: {mail_content}') mail.close() mail.logout() print(sender_email) df = pd.DataFrame() df['Contact Name'] = sender_names df['Email Address'] = sender_email df['Company'] = Sender_company df.to_csv('contacts.csv',encoding='utf-8') print("Searching has been completed")
我试过用带x计数器的if循环缩小列表规模,但没解决问题,求排查原因。
原因分析与修复方案
核心问题
你在创建DataFrame时犯了两个明显错误:
- 把单个变量
sender_email赋值给了Email Address列,而不是存储所有邮箱地址的列表sender_list。你已经把符合条件的邮箱都添加到了sender_list,但最后构建DataFrame时用了循环最后一次迭代的sender_email值,导致整列都是同一个邮箱。 y+1只是计算了值但没有赋值回y,导致后续生成的CSV文件名不会递增。
修复后的代码
import email import imaplib import re import pandas as pd EMAIL = "email" PASSWORD = 'password' SERVER = "imap.example.com" sender_list = [] sender_names = [] Sender_company = [] Sender_phone = {} x = 0 y = 1 # connect to the server and go to its inbox mail = imaplib.IMAP4_SSL(SERVER) sock=mail.socket() timeout = 60 * 20 # 5 minutes sock.settimeout(timeout) mail.login(EMAIL, PASSWORD) # we choose the inbox but you can select others mail.select('inbox') # we'll search using the ALL criteria to retrieve # every message inside the inbox # it will return with its status and a list of ids status, data = mail.search(None, 'ALL') # the list returned is a list of bytes separated # by white spaces on this format: [b'1 2 3', b'4 5 6'] # so, to separate it first we create an empty list mail_ids = [] # then we go through the list splitting its blocks # of bytes and appending to the mail_ids list for block in data: # the split function called without parameter # transforms the text or bytes into a list using # as separator the white spaces: # b'1 2 3'.split() => [b'1', b'2', b'3'] mail_ids += block.split() # now for every id we'll fetch the email # to extract its content for i in mail_ids: # the fetch function fetch the email given its id # and format that you want the message to be status, data = mail.fetch(i, '(RFC822)') # the content data at the '(RFC822)' format comes on # a list with a tuple with header, content, and the closing # byte b')' for response_part in data: # so if its a tuple... if isinstance(response_part, tuple): # we go for the content at its second element # skipping the header at the first and the closing # at the third message = email.message_from_bytes(response_part[1]) # with the content we can extract the info about # who sent the message and its subject mail_from = message['from'] mail_subject = message['subject'] sender_formatted = mail_from.replace(">","").split("<") list_length = len(sender_formatted) if list_length != 2: continue sender_email = sender_formatted[1] sender_name = sender_formatted[0] if sender_list.count(sender_email) > 0 : continue # then for the text we have a little more work to do # because it can be in plain text or multipart # if its not plain text we need to separate the message # from its annexes to get the text if message.is_multipart(): mail_content = '' # on multipart we have the text message and # another things like annex, and html version # of the message, in that case we loop through # the email payload for part in message.get_payload(): # if the content type is text/plain # we extract it if part.get_content_type() == 'text/plain': mail_content += part.get_payload() else: # if the message isn't multipart, just extract it mail_content = message.get_payload() words=["MF116","MF115","install","removal","U6","U16","U25","U40","U65","U100","meter","installed"] content= mail_content # print(content) for word in words: result = re.search(word,content) if result : # print (f"search success: {word}") sender_list.append(sender_email) sender_names.append(sender_name) company = sender_email.split("@") company = company[1].split(".") company = company[0] Sender_company.append(company) x += 1 if x == 100 : df = pd.DataFrame() df['Contact Name'] = sender_names df['Email Address'] = sender_list # 改用存储所有邮箱的列表 df['Company'] = Sender_company file_name = (f'contacts{y}.csv') df.to_csv(file_name,encoding='utf-8') y += 1 # 正确赋值让序号递增 # print(sender_email,company, sender_name) break else : # print(f"Search failed: {word}") continue # and then let's show its result # print(f'From: {mail_from}') # print(f'Subject: {mail_subject}') # print(f'Content: {mail_content}') mail.close() mail.logout() print(sender_email) df = pd.DataFrame() df['Contact Name'] = sender_names df['Email Address'] = sender_list # 同样改用存储所有邮箱的列表 df['Company'] = Sender_company df.to_csv('contacts.csv',encoding='utf-8') print("Searching has been completed")
关键修改点
- 修正邮箱列赋值:两处构建DataFrame的地方,将
df['Email Address'] = sender_email改为df['Email Address'] = sender_list,用存储所有符合条件邮箱的列表替换单个变量。 - 修正计数器递增:将
y+1改为y += 1,确保每次生成CSV后文件名的序号正确递增。
内容的提问来源于stack exchange,提问作者gm69755
相关产品推荐
相关产品推荐

