You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何捕获DNS_PROBE_FINISHED_NXDOMAIN异常,避免Jupyter Notebook卡顿?

解决DNS_PROBE_FINISHED_NXDOMAIN异常捕获与程序卡顿问题

问题描述

从URL数据集中提取域名并通过whois数据库获取信息时,Jupyter Notebook程序出现卡顿,同时遇到DNS解析失败错误:

This site can’t be reached
Check if there is a typo in www.content.usatoday.com.
If spelling is correct, try running Windows Network Diagnostics.
DNS_PROBE_FINISHED_NXDOMAIN

现有核心代码如下:

函数定义

def get_features(url,label):
    
    features = []
    ....
    
    dnsrecord = 0
    try:
        domain = whois.whois(urlparse(url).netloc)
    except:
        dnsrecord = 1 
     
    features.append(dnsrecord)
    
    ....
    return features 

运行代码

features = []
    
for i in range (0,len(df)):
    url = df['url'][i]
    label = df['label'][i]
    features.append(feature_extraction(url,label))

问题分析

  1. 现有try-except捕获所有异常,无法精准识别DNS解析失败的场景;
  2. whois查询默认无超时限制,遇到无效域名时会持续等待,导致程序卡顿;
  3. 循环遍历方式效率较低,易加剧卡顿问题。

解决方案

1. 精准捕获DNS解析异常

DNS_PROBE_FINISHED_NXDOMAIN对应底层的socket.gaierror异常,可针对性捕获;同时捕获whois库专属异常及超时异常,区分不同错误类型。

2. 添加查询超时机制

通过装饰器限制whois查询的最大等待时间,避免单个域名查询卡住整个程序。

3. 优化循环遍历方式

改用DataFrame的iterrows()迭代,代码更简洁高效。

修改后的完整代码

import socket
from urllib.parse import urlparse
import whois
from functools import wraps
import time

# 超时限制装饰器
def timeout(seconds=10):
    def decorator(func):
        @wraps(func)
        def wrapper(*args, **kwargs):
            result = None
            def target():
                nonlocal result
                result = func(*args, **kwargs)
            import threading
            thread = threading.Thread(target=target)
            thread.daemon = True
            thread.start()
            thread.join(seconds)
            if thread.is_alive():
                raise TimeoutError(f"查询超时,已超过{seconds}秒")
            return result
        return wrapper
    return decorator

def get_features(url, label):
    features = []
    # 其他特征提取逻辑...
    
    dnsrecord = 0
    domain = urlparse(url).netloc
    
    # 先判断域名是否有效
    if not domain:
        dnsrecord = 1
        features.append(dnsrecord)
        # 其他逻辑...
        return features
    
    # 带超时的whois查询
    @timeout(5)  # 设置5秒超时,可根据需求调整
    def query_whois(domain):
        return whois.whois(domain)
    
    try:
        whois_info = query_whois(domain)
        # 这里可以添加whois信息的特征提取逻辑
    except socket.gaierror:
        # 捕获DNS解析失败(对应DNS_PROBE_FINISHED_NXDOMAIN)
        dnsrecord = 1
    except TimeoutError:
        # 捕获查询超时
        dnsrecord = 2
    except whois.parser.PywhoisError:
        # 捕获whois库返回的解析错误
        dnsrecord = 3
    except Exception as e:
        # 兜底处理其他意外错误
        dnsrecord = 4
        print(f"域名{domain}处理出错:{str(e)}")
    
    features.append(dnsrecord)
    # 其他特征提取逻辑...
    return features

# 优化后的循环调用
features = []
for idx, row in df.iterrows():
    url = row['url']
    label = row['label']
    features.append(get_features(url, label))

说明

  • socket.gaierror直接对应DNS解析失败的场景,精准捕获即可处理DNS_PROBE_FINISHED_NXDOMAIN错误;
  • 超时装饰器通过多线程实现,避免阻塞主进程;
  • 用不同的dnsrecord值标记不同错误类型,便于后续数据分析排查问题。

内容的提问来源于stack exchange,提问作者wong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 17:01:06