You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python邮件脚本故障:Email列表所有项为同一值

问题

我写了一个脚本遍历邮件收件箱,找出含指定关键词的邮件,把发件人姓名、邮箱、公司信息分别存入列表,最后导入Pandas DataFrame并导出为CSV。现在姓名和公司列都正常,但邮箱地址列所有项都是同一个值,无法正确存储不同邮箱。

预期邮箱列表:

123@ABC.com 
234@boo.com 
info@example.com 

实际得到的邮箱列表:

123@ABC.com 
123@ABC.com 
123@ABC.com 

完整代码:

import email
import imaplib
import re
import pandas as pd
EMAIL = "email"
PASSWORD = 'password'
SERVER = "imap.example.com"

sender_list = []
sender_names = []
Sender_company = []
Sender_phone = {}
x = 0 
y = 1
# connect to the server and go to its inbox
mail = imaplib.IMAP4_SSL(SERVER)
sock=mail.socket()
timeout = 60 * 20 # 5 minutes
sock.settimeout(timeout)
mail.login(EMAIL, PASSWORD)
# we choose the inbox but you can select others
mail.select('inbox')

# we'll search using the ALL criteria to retrieve
# every message inside the inbox
# it will return with its status and a list of ids
status, data = mail.search(None, 'ALL')
# the list returned is a list of bytes separated
# by white spaces on this format: [b'1 2 3', b'4 5 6']
# so, to separate it first we create an empty list
mail_ids = []
# then we go through the list splitting its blocks
# of bytes and appending to the mail_ids list
for block in data:
    # the split function called without parameter
    # transforms the text or bytes into a list using
    # as separator the white spaces:
    # b'1 2 3'.split() => [b'1', b'2', b'3']
    mail_ids += block.split()

# now for every id we'll fetch the email
# to extract its content
for i in mail_ids:
    # the fetch function fetch the email given its id
    # and format that you want the message to be
    status, data = mail.fetch(i, '(RFC822)')

    # the content data at the '(RFC822)' format comes on
    # a list with a tuple with header, content, and the closing
    # byte b')'
    for response_part in data:
        # so if its a tuple...
        if isinstance(response_part, tuple):
            # we go for the content at its second element
            # skipping the header at the first and the closing
            # at the third
            message = email.message_from_bytes(response_part[1])

            # with the content we can extract the info about
            # who sent the message and its subject
            mail_from = message['from']
            mail_subject = message['subject']
            sender_formatted = mail_from.replace(">","").split("<")
            list_length = len(sender_formatted)
            if list_length != 2: 
                continue
            sender_email = sender_formatted[1]
            sender_name = sender_formatted[0]
            if sender_list.count(sender_email) > 0 : 
                continue

            # then for the text we have a little more work to do
            # because it can be in plain text or multipart
            # if its not plain text we need to separate the message
            # from its annexes to get the text
            if message.is_multipart():
                mail_content = ''

                # on multipart we have the text message and
                # another things like annex, and html version
                # of the message, in that case we loop through
                # the email payload
                for part in message.get_payload():
                    # if the content type is text/plain
                    # we extract it
                    if part.get_content_type() == 'text/plain':
                        mail_content += part.get_payload()
            else:
                # if the message isn't multipart, just extract it
                mail_content = message.get_payload()
            words=["MF116","MF115","install","removal","U6","U16","U25","U40","U65","U100","meter","installed"]
            content= mail_content 
            # print(content)
            for word in words:
                result = re.search(word,content)
                if result : 
                    # print (f"search success: {word}")
                    sender_list.append(sender_email)
                    sender_names.append(sender_name)

                    company = sender_email.split("@")
                    company = company[1].split(".")
                    company = company[0]
                    Sender_company.append(company)
                    x += 1 
                    if x == 100 :
                        df = pd.DataFrame()
                        df['Contact Name'] = sender_names
                        df['Email Address'] = sender_email
                        df['Company'] = Sender_company
                        file_name = (f'contacts{y}.csv')
                        df.to_csv(file_name,encoding='utf-8')
                        y+1 

                
                    # print(sender_email,company, sender_name)
                    break
                else : 
                    # print(f"Search failed: {word}")
                    continue

            # and then let's show its result
            # print(f'From: {mail_from}')
            # print(f'Subject: {mail_subject}')
            # print(f'Content: {mail_content}')
            


mail.close()
mail.logout()
print(sender_email)
df = pd.DataFrame()
df['Contact Name'] = sender_names
df['Email Address'] = sender_email
df['Company'] = Sender_company

df.to_csv('contacts.csv',encoding='utf-8')

print("Searching has been completed")

我试过用带x计数器的if循环缩小列表规模,但没解决问题,求排查原因。


原因分析与修复方案

核心问题

你在创建DataFrame时犯了两个明显错误:

  1. 把单个变量sender_email赋值给了Email Address列,而不是存储所有邮箱地址的列表sender_list。你已经把符合条件的邮箱都添加到了sender_list,但最后构建DataFrame时用了循环最后一次迭代的sender_email值,导致整列都是同一个邮箱。
  2. y+1只是计算了值但没有赋值回y,导致后续生成的CSV文件名不会递增。

修复后的代码

import email
import imaplib
import re
import pandas as pd
EMAIL = "email"
PASSWORD = 'password'
SERVER = "imap.example.com"

sender_list = []
sender_names = []
Sender_company = []
Sender_phone = {}
x = 0 
y = 1
# connect to the server and go to its inbox
mail = imaplib.IMAP4_SSL(SERVER)
sock=mail.socket()
timeout = 60 * 20 # 5 minutes
sock.settimeout(timeout)
mail.login(EMAIL, PASSWORD)
# we choose the inbox but you can select others
mail.select('inbox')

# we'll search using the ALL criteria to retrieve
# every message inside the inbox
# it will return with its status and a list of ids
status, data = mail.search(None, 'ALL')
# the list returned is a list of bytes separated
# by white spaces on this format: [b'1 2 3', b'4 5 6']
# so, to separate it first we create an empty list
mail_ids = []
# then we go through the list splitting its blocks
# of bytes and appending to the mail_ids list
for block in data:
    # the split function called without parameter
    # transforms the text or bytes into a list using
    # as separator the white spaces:
    # b'1 2 3'.split() => [b'1', b'2', b'3']
    mail_ids += block.split()

# now for every id we'll fetch the email
# to extract its content
for i in mail_ids:
    # the fetch function fetch the email given its id
    # and format that you want the message to be
    status, data = mail.fetch(i, '(RFC822)')

    # the content data at the '(RFC822)' format comes on
    # a list with a tuple with header, content, and the closing
    # byte b')'
    for response_part in data:
        # so if its a tuple...
        if isinstance(response_part, tuple):
            # we go for the content at its second element
            # skipping the header at the first and the closing
            # at the third
            message = email.message_from_bytes(response_part[1])

            # with the content we can extract the info about
            # who sent the message and its subject
            mail_from = message['from']
            mail_subject = message['subject']
            sender_formatted = mail_from.replace(">","").split("<")
            list_length = len(sender_formatted)
            if list_length != 2: 
                continue
            sender_email = sender_formatted[1]
            sender_name = sender_formatted[0]
            if sender_list.count(sender_email) > 0 : 
                continue

            # then for the text we have a little more work to do
            # because it can be in plain text or multipart
            # if its not plain text we need to separate the message
            # from its annexes to get the text
            if message.is_multipart():
                mail_content = ''

                # on multipart we have the text message and
                # another things like annex, and html version
                # of the message, in that case we loop through
                # the email payload
                for part in message.get_payload():
                    # if the content type is text/plain
                    # we extract it
                    if part.get_content_type() == 'text/plain':
                        mail_content += part.get_payload()
            else:
                # if the message isn't multipart, just extract it
                mail_content = message.get_payload()
            words=["MF116","MF115","install","removal","U6","U16","U25","U40","U65","U100","meter","installed"]
            content= mail_content 
            # print(content)
            for word in words:
                result = re.search(word,content)
                if result : 
                    # print (f"search success: {word}")
                    sender_list.append(sender_email)
                    sender_names.append(sender_name)

                    company = sender_email.split("@")
                    company = company[1].split(".")
                    company = company[0]
                    Sender_company.append(company)
                    x += 1 
                    if x == 100 :
                        df = pd.DataFrame()
                        df['Contact Name'] = sender_names
                        df['Email Address'] = sender_list  # 改用存储所有邮箱的列表
                        df['Company'] = Sender_company
                        file_name = (f'contacts{y}.csv')
                        df.to_csv(file_name,encoding='utf-8')
                        y += 1  # 正确赋值让序号递增

                
                    # print(sender_email,company, sender_name)
                    break
                else : 
                    # print(f"Search failed: {word}")
                    continue

            # and then let's show its result
            # print(f'From: {mail_from}')
            # print(f'Subject: {mail_subject}')
            # print(f'Content: {mail_content}')
            


mail.close()
mail.logout()
print(sender_email)
df = pd.DataFrame()
df['Contact Name'] = sender_names
df['Email Address'] = sender_list  # 同样改用存储所有邮箱的列表
df['Company'] = Sender_company

df.to_csv('contacts.csv',encoding='utf-8')

print("Searching has been completed")

关键修改点

  1. 修正邮箱列赋值:两处构建DataFrame的地方,将df['Email Address'] = sender_email改为df['Email Address'] = sender_list,用存储所有符合条件邮箱的列表替换单个变量。
  2. 修正计数器递增:将y+1改为y += 1,确保每次生成CSV后文件名的序号正确递增。

内容的提问来源于stack exchange,提问作者gm69755

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 02:55:24