You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

python-whois模块返回结果不稳定,同一域名输出类型不一致求助

解决python-whois返回邮箱类型不一致及多次查询结果差异问题

问题原因分析

你遇到的问题主要来自两个方面:

  1. python-whois模块特性:该模块依赖whois服务器返回的原始数据,不同服务器或同一服务器不同时间返回的数据格式可能存在波动,模块解析时未做强制统一的类型转换,导致同一域名的emails/registrant_email字段有时返回字符串(单邮箱场景)、有时返回列表(多邮箱场景)。
  2. 重复查询放大差异:原代码中多次调用whois.whois(s[0]),每次调用都会重新发起网络请求,不同请求的目标服务器、返回数据格式可能不同,进一步加剧了结果的不一致性。

优化解决方案

通过「单次查询复用结果+统一类型转换」彻底解决问题,优化后的代码如下:

import whois
import csv

def normalize_email(data):
    """统一邮箱格式,确保返回纯字符串组成的列表"""
    if data is None:
        return []
    # 单邮箱字符串转为单元素列表
    if isinstance(data, str):
        return [data.strip()]
    # 多邮箱列表处理,同时清理嵌套结构(部分返回会带额外信息的元组)
    elif isinstance(data, list):
        cleaned = []
        for item in data:
            if isinstance(item, str):
                cleaned.append(item.strip())
            else:
                # 提取元组/嵌套列表中的邮箱字符串
                cleaned.append(str(item[0]).strip())
        return cleaned
    return []

with open('emails.csv', 'r', newline='') as f:
    reader = csv.reader(f)
    distro = [row for row in reader]

with open('emails_done.csv', 'w', newline='') as f:
    wr = csv.writer(f)
    for s in distro:
        domain = s[0]
        try:
            # 每个域名仅查询一次whois数据,复用结果
            whois_result = whois.whois(domain)
            
            # 优先处理emails字段
            emails = normalize_email(whois_result.emails)
            if emails:
                wr.writerow([domain] + emails)
                continue
            
            # 处理registrant_email字段
            reg_emails = normalize_email(whois_result.registrant_email)
            if reg_emails:
                wr.writerow([domain] + reg_emails)
                continue
            
            # 无匹配邮箱的情况
            wr.writerow([domain, 'nothing found'])
        except Exception as e:
            print(f"域名{domain}处理失败: {str(e)}")
            wr.writerow([domain, 'invaliddomain'])

关键改动说明

  • 单次查询复用:每个域名仅调用一次whois.whois,避免重复网络请求带来的结果波动,同时提升脚本运行效率。
  • 类型统一转换:normalize_email函数将所有邮箱相关字段统一转为纯字符串列表,彻底解决csv.writerow的格式兼容问题,同时清理了部分返回结果中的嵌套冗余信息。
  • 简化逻辑:按优先级依次检查emails、registrant_email字段,代码逻辑更清晰,减少冗余判断。

内容的提问来源于stack exchange,提问作者Xanya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 03:45:23