You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多进程Office365邮箱校验时全局数组无法追加内容如何解决

问题核心原因
  • Python多进程不存在共享的全局内存空间,子进程对全局变量的修改仅作用于自身内存副本,不会同步到主进程,因此主进程读取的GOOD_EMAILS_ARRAY、BAD_EMAILS_ARRAY始终为初始空值
  • 附加问题:newChecker函数的有效性判断逻辑冗余重复,进程池没有等待所有子进程执行完成就调用写文件逻辑、全局数组未预先初始化、进程池开500个进程远超常规服务器承载上限,这些问题也会导致异常
修复方案

推荐使用「子进程返回校验结果,主进程统一收集写入」的方案,规避共享变量的同步问题,修改后代码如下:

import os
import re
import requests as req
import multiprocessing
from datetime import datetime

currentDirectory = os.getcwd()  # set the current directory - /new/

# 存储路径配置
location_emails_goods = currentDirectory + '/contacts/goods/'
location_emails_bads = currentDirectory + '/contacts/bads/'
location_emails = currentDirectory + '/contacts/contacts.txt'

now = datetime.now()
todayString = now.strftime('%d-%m-%Y-%H-%M-%S')
url = 'https://login.microsoftonline.com/common/GetCredentialType'

# 提前创建目录避免写入报错
os.makedirs(location_emails_goods, exist_ok=True)
os.makedirs(location_emails_bads, exist_ok=True)

# 读取待校验邮箱列表
def get_contacts(filename):
    emails = []
    with open(filename, mode='r', encoding='utf-8') as contacts_file:
        for a_contact in contacts_file:
            stripped = a_contact.strip()
            if stripped:
                emails.append(stripped)
    return emails

# 校验函数直接返回结果,不操作全局变量
def newChecker(email):
    s = req.session()
    body = '{"Username":"%s"}' % email
    try:
        request = req.post(url, data=body, timeout=10)
        response = request.text
        if '"IfExistsResult":0,' in response:
            return (email, 'good')
        elif '"IfExistsResult":1,' in response:
            return (email, 'bad')
        else:
            return (email, 'bad')
    except Exception:
        return (email, 'bad')

def mp_handler(p, all_emails):
    return p.map(newChecker, all_emails)

if __name__ == '__main__':
    ALL_EMAILS = get_contacts(location_emails)
    # 进程数设置为CPU核心数的2倍即可,过高会导致系统资源耗尽
    p = multiprocessing.Pool(multiprocessing.cpu_count() * 2)
    results = mp_handler(p, ALL_EMAILS)
    p.close()
    p.join() # 等待所有子进程执行完成再继续后续操作

    # 主进程统一写入结果
    FILE_NAME_DATE_GOODS = f"{location_emails_goods}{todayString}.txt"
    FILE_NAME_DATE_BADS = f"{location_emails_bads}{todayString}.txt"
    with open(FILE_NAME_DATE_GOODS, 'w', encoding='utf-8') as f_good, open(FILE_NAME_DATE_BADS, 'w', encoding='utf-8') as f_bad:
        for email, status in results:
            if status == 'good':
                f_good.write(email + '\n')
            else:
                f_bad.write(email + '\n')
关键修改点
  • 去掉全局变量共享逻辑,校验函数直接返回邮箱和对应的校验状态,主进程统一收集所有结果
  • 增加目录提前创建、网络请求异常捕获、超时设置,避免边界场景报错
  • 调整进程池数量为CPU核心数的2倍,避免过多进程导致系统卡顿崩溃
  • 调整文件写入逻辑为所有进程执行完成后统一用with上下文管理操作,避免文件句柄泄漏
  • 优化有效性判断逻辑,用字符串匹配代替正则,性能更高且避免正则匹配失误

内容的提问来源于stack exchange,提问作者rident smith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 08:45:01