如何捕获DNS_PROBE_FINISHED_NXDOMAIN异常,避免Jupyter Notebook卡顿?
解决DNS_PROBE_FINISHED_NXDOMAIN异常捕获与程序卡顿问题
问题描述
从URL数据集中提取域名并通过whois数据库获取信息时,Jupyter Notebook程序出现卡顿,同时遇到DNS解析失败错误:
This site can’t be reached
Check if there is a typo in www.content.usatoday.com.
If spelling is correct, try running Windows Network Diagnostics.
DNS_PROBE_FINISHED_NXDOMAIN
现有核心代码如下:
函数定义
def get_features(url,label): features = [] .... dnsrecord = 0 try: domain = whois.whois(urlparse(url).netloc) except: dnsrecord = 1 features.append(dnsrecord) .... return features
运行代码
features = [] for i in range (0,len(df)): url = df['url'][i] label = df['label'][i] features.append(feature_extraction(url,label))
问题分析
- 现有
try-except捕获所有异常,无法精准识别DNS解析失败的场景; - whois查询默认无超时限制,遇到无效域名时会持续等待,导致程序卡顿;
- 循环遍历方式效率较低,易加剧卡顿问题。
解决方案
1. 精准捕获DNS解析异常
DNS_PROBE_FINISHED_NXDOMAIN对应底层的socket.gaierror异常,可针对性捕获;同时捕获whois库专属异常及超时异常,区分不同错误类型。
2. 添加查询超时机制
通过装饰器限制whois查询的最大等待时间,避免单个域名查询卡住整个程序。
3. 优化循环遍历方式
改用DataFrame的iterrows()迭代,代码更简洁高效。
修改后的完整代码
import socket from urllib.parse import urlparse import whois from functools import wraps import time # 超时限制装饰器 def timeout(seconds=10): def decorator(func): @wraps(func) def wrapper(*args, **kwargs): result = None def target(): nonlocal result result = func(*args, **kwargs) import threading thread = threading.Thread(target=target) thread.daemon = True thread.start() thread.join(seconds) if thread.is_alive(): raise TimeoutError(f"查询超时,已超过{seconds}秒") return result return wrapper return decorator def get_features(url, label): features = [] # 其他特征提取逻辑... dnsrecord = 0 domain = urlparse(url).netloc # 先判断域名是否有效 if not domain: dnsrecord = 1 features.append(dnsrecord) # 其他逻辑... return features # 带超时的whois查询 @timeout(5) # 设置5秒超时,可根据需求调整 def query_whois(domain): return whois.whois(domain) try: whois_info = query_whois(domain) # 这里可以添加whois信息的特征提取逻辑 except socket.gaierror: # 捕获DNS解析失败(对应DNS_PROBE_FINISHED_NXDOMAIN) dnsrecord = 1 except TimeoutError: # 捕获查询超时 dnsrecord = 2 except whois.parser.PywhoisError: # 捕获whois库返回的解析错误 dnsrecord = 3 except Exception as e: # 兜底处理其他意外错误 dnsrecord = 4 print(f"域名{domain}处理出错:{str(e)}") features.append(dnsrecord) # 其他特征提取逻辑... return features # 优化后的循环调用 features = [] for idx, row in df.iterrows(): url = row['url'] label = row['label'] features.append(get_features(url, label))
说明
socket.gaierror直接对应DNS解析失败的场景,精准捕获即可处理DNS_PROBE_FINISHED_NXDOMAIN错误;- 超时装饰器通过多线程实现,避免阻塞主进程;
- 用不同的
dnsrecord值标记不同错误类型,便于后续数据分析排查问题。
内容的提问来源于stack exchange,提问作者wong
相关产品推荐
相关产品推荐

