使用NetworkX与Requests获取部分网站URL的IP地址失败问题排查
问题背景
尝试将网站URL映射到对应IP地址,并用NetworkX绘制关系图。通过BeautifulSoup提取网页链接进行BFS遍历,曾尝试直接为节点分配属性,后改用set_node_attributes结合URL和IP列表设置属性,但https://www.youtube.com/about/#content始终无法获取IP属性,后续代码读取该节点ip_address时触发KeyError。
用户代码
获取链接与IP的函数
# Gets the links to traverse as well as the IP Address and returns them. def get_links(url): response = requests.get(url, stream=True) ip_address = response.raw._connection.sock.getsockname() soup = BeautifulSoup(response.text, 'lxml') return [urljoin(url, a['href']) for a in soup.find_all('a', href=True)], ip_address
BFS遍历函数
# Traverses the urls and def bfs_traversal(start_url, num_layers, graph): visited = [] visited_ip_addresses = [] queue = deque([(start_url, 0)]) attrs = {} while queue: url, layer = queue.popleft() if layer > num_layers: # add our attrs for url, ip_address in zip(visited, visited_ip_addresses): attrs[url] = ip_address # Need to set the node attributes manually see below nx.set_node_attributes(graph, values=attrs, name='ip_address') break if url not in visited: visited.append(url) connections, ip_address = get_links(url) visited_ip_addresses.append(ip_address) # No longer seems to work hence the above code graph.add_node(url, ip_address=ip_address) # throws error graph.add_node(url, {'ip_address':ip_address}) for x in connections: graph.add_edge(url, x) queue.append((x, layer + 1))
图初始化与位置设置
scan_layers = 1 entrance = "https://www.youtube.com/" graph = nx.Graph() seed = 0 bfs_traversal(entrance, scan_layers, graph) pos = nx.spring_layout(graph, seed=seed) # Add pos to the node for n, p in pos.items(): graph.nodes[n]['pos'] = p
触发错误的代码
ip_addresses = [] for node in graph.nodes(data=True): print(node, '\n') ip_address = node[1]['ip_address'][0] ip_addresses.append(ip_address)
报错信息
('https://www.youtube.com/about/#content', {'pos': array([-0.48220751, -0.4565694 ])}) Traceback (most recent call last): File "C:\Users\Owner\PycharmProjects\WebsiteMapper\gui.py", line 125, in <module> gp = GraphPage() ^^^^^^^^^^^ File "C:\Users\Owner\PycharmProjects\WebsiteMapper\gui.py", line 36, in __init__ fig_json = create_network_graph() ^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\Owner\PycharmProjects\WebsiteMapper\graphs.py", line 89, in create_network_graph ip_address = node[1]['ip_address'][0] ~~~~~~~^^^^^^^^^^^^^^ KeyError: 'ip_address'
问题原因
网络请求未捕获异常:
get_links函数直接调用requests.get,若请求带锚点的URL时出现网络超时、HTTP错误(如403、404),会抛出异常中断流程。此时visited.append(url)已执行,但visited_ip_addresses.append(ip_address)和节点属性设置步骤未执行,导致该URL在visited中却无对应IP,后续批量设置属性时被遗漏。BFS属性设置时机错误:仅当
layer > num_layers时才执行set_node_attributes,若队列提前处理完毕(无层级超过设定值的节点),所有节点IP属性都不会被设置。同时add_node已设置属性,后续批量设置属于冗余操作。锚点URL的请求特性:带
#content的锚点URL,服务器通常返回与主页面相同内容,但requests.get可能因重定向、服务器反爬策略导致请求失败,进而中断IP获取流程。
修复方案
1. 为网络请求添加异常处理
确保请求失败时仍能返回默认值,避免流程中断:
def get_links(url): try: response = requests.get(url, stream=True, timeout=5) response.raise_for_status() # 触发HTTP错误异常 ip_address = response.raw._connection.sock.getsockname() soup = BeautifulSoup(response.text, 'lxml') links = [urljoin(url, a['href']) for a in soup.find_all('a', href=True)] return links, ip_address except Exception as e: print(f"请求URL失败 {url}: {str(e)}") return [], None # 返回空链接和None标记IP获取失败
2. 调整BFS逻辑,实时设置节点属性
移除冗余的批量设置逻辑,处理每个节点时直接设置属性,同时优化visited为集合提升效率:
def bfs_traversal(start_url, num_layers, graph): visited = set() # 集合查找效率远高于列表 queue = deque([(start_url, 0)]) while queue: url, layer = queue.popleft() # 跳过超过层级或已访问的节点 if layer > num_layers or url in visited: continue visited.add(url) connections, ip_address = get_links(url) # 直接设置节点属性,兼容节点已存在的情况 graph.nodes[url]['ip_address'] = ip_address if ip_address is not None else ("无法获取",) for x in connections: graph.add_edge(url, x) if x not in visited: queue.append((x, layer + 1))
3. 容错处理IP属性读取
在读取IP时添加KeyError和空值判断,避免程序崩溃:
ip_addresses = [] for node in graph.nodes(data=True): print(node, '\n') # 容错:获取ip_address属性,不存在则返回默认值 ip_info = node[1].get('ip_address', ("无法获取",)) ip_address = ip_info[0] ip_addresses.append(ip_address)
4. 可选:归一化锚点URL(减少重复节点)
如果不需要区分带锚点和不带锚点的URL,可移除锚点部分:
from urllib.parse import urlparse, urlunparse def normalize_url(url): parsed = urlparse(url) # 移除URL中的锚点片段 return urlunparse(parsed._replace(fragment='')) # 在get_links中替换链接处理逻辑 links = [normalize_url(urljoin(url, a['href'])) for a in soup.find_all('a', href=True)]
额外优化建议
- 使用
set存储visited,避免O(n)时间复杂度的查找操作。 - 为
requests.get添加timeout参数,防止请求长时间阻塞。 - 可添加User-Agent请求头,避免被服务器识别为爬虫导致请求失败。
内容的提问来源于stack exchange,提问作者David Frick

