You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用NetworkX与Requests获取部分网站URL的IP地址失败问题排查

URL节点缺失IP属性的问题分析与修复

问题背景

尝试将网站URL映射到对应IP地址,并用NetworkX绘制关系图。通过BeautifulSoup提取网页链接进行BFS遍历,曾尝试直接为节点分配属性,后改用set_node_attributes结合URL和IP列表设置属性,但https://www.youtube.com/about/#content始终无法获取IP属性,后续代码读取该节点ip_address时触发KeyError。

用户代码

获取链接与IP的函数

# Gets the links to traverse as well as the IP Address and returns them.
def get_links(url):
    response = requests.get(url, stream=True)
    ip_address = response.raw._connection.sock.getsockname()
    soup = BeautifulSoup(response.text, 'lxml')
    return [urljoin(url, a['href']) for a in soup.find_all('a', href=True)], ip_address

BFS遍历函数

# Traverses the urls and 
def bfs_traversal(start_url, num_layers, graph):
    visited = []
    visited_ip_addresses = []
    queue = deque([(start_url, 0)])

    attrs = {}
    while queue:
        url, layer = queue.popleft()
        if layer > num_layers:
            # add our attrs
            for url, ip_address in zip(visited, visited_ip_addresses):
                attrs[url] = ip_address

            # Need to set the node attributes manually see below
            nx.set_node_attributes(graph, values=attrs, name='ip_address')

            break

        if url not in visited:
            visited.append(url)
            connections, ip_address = get_links(url)
            visited_ip_addresses.append(ip_address)

            # No longer seems to work hence the above code
            graph.add_node(url, ip_address=ip_address)
            # throws error graph.add_node(url, {'ip_address':ip_address})


            for x in connections:
                graph.add_edge(url, x)
                queue.append((x, layer + 1))

图初始化与位置设置

scan_layers = 1
entrance = "https://www.youtube.com/"

graph = nx.Graph()
seed = 0

bfs_traversal(entrance, scan_layers, graph)
pos = nx.spring_layout(graph, seed=seed)

# Add pos to the node
for n, p in pos.items():
    graph.nodes[n]['pos'] = p

触发错误的代码

ip_addresses = []
for node in graph.nodes(data=True):
    print(node,  '\n')
    ip_address = node[1]['ip_address'][0]
    ip_addresses.append(ip_address)

报错信息

('https://www.youtube.com/about/#content', {'pos': array([-0.48220751, -0.4565694 ])})

Traceback (most recent call last):
  File "C:\Users\Owner\PycharmProjects\WebsiteMapper\gui.py", line 125, in <module>
    gp = GraphPage()
         ^^^^^^^^^^^
  File "C:\Users\Owner\PycharmProjects\WebsiteMapper\gui.py", line 36, in __init__
    fig_json = create_network_graph()
               ^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\Owner\PycharmProjects\WebsiteMapper\graphs.py", line 89, in create_network_graph
    ip_address = node[1]['ip_address'][0]
                 ~~~~~~~^^^^^^^^^^^^^^
KeyError: 'ip_address'

问题原因

  1. 网络请求未捕获异常:get_links函数直接调用requests.get,若请求带锚点的URL时出现网络超时、HTTP错误(如403、404),会抛出异常中断流程。此时visited.append(url)已执行,但visited_ip_addresses.append(ip_address)和节点属性设置步骤未执行,导致该URL在visited中却无对应IP,后续批量设置属性时被遗漏。

  2. BFS属性设置时机错误:仅当layer > num_layers时才执行set_node_attributes,若队列提前处理完毕(无层级超过设定值的节点),所有节点IP属性都不会被设置。同时add_node已设置属性,后续批量设置属于冗余操作。

  3. 锚点URL的请求特性:带#content的锚点URL,服务器通常返回与主页面相同内容,但requests.get可能因重定向、服务器反爬策略导致请求失败,进而中断IP获取流程。

修复方案

1. 为网络请求添加异常处理

确保请求失败时仍能返回默认值,避免流程中断:

def get_links(url):
    try:
        response = requests.get(url, stream=True, timeout=5)
        response.raise_for_status()  # 触发HTTP错误异常
        ip_address = response.raw._connection.sock.getsockname()
        soup = BeautifulSoup(response.text, 'lxml')
        links = [urljoin(url, a['href']) for a in soup.find_all('a', href=True)]
        return links, ip_address
    except Exception as e:
        print(f"请求URL失败 {url}: {str(e)}")
        return [], None  # 返回空链接和None标记IP获取失败

2. 调整BFS逻辑,实时设置节点属性

移除冗余的批量设置逻辑,处理每个节点时直接设置属性,同时优化visited为集合提升效率:

def bfs_traversal(start_url, num_layers, graph):
    visited = set()  # 集合查找效率远高于列表
    queue = deque([(start_url, 0)])

    while queue:
        url, layer = queue.popleft()
        # 跳过超过层级或已访问的节点
        if layer > num_layers or url in visited:
            continue

        visited.add(url)
        connections, ip_address = get_links(url)
        
        # 直接设置节点属性,兼容节点已存在的情况
        graph.nodes[url]['ip_address'] = ip_address if ip_address is not None else ("无法获取",)

        for x in connections:
            graph.add_edge(url, x)
            if x not in visited:
                queue.append((x, layer + 1))

3. 容错处理IP属性读取

在读取IP时添加KeyError和空值判断,避免程序崩溃:

ip_addresses = []
for node in graph.nodes(data=True):
    print(node, '\n')
    # 容错:获取ip_address属性,不存在则返回默认值
    ip_info = node[1].get('ip_address', ("无法获取",))
    ip_address = ip_info[0]
    ip_addresses.append(ip_address)

4. 可选:归一化锚点URL(减少重复节点)

如果不需要区分带锚点和不带锚点的URL,可移除锚点部分:

from urllib.parse import urlparse, urlunparse

def normalize_url(url):
    parsed = urlparse(url)
    # 移除URL中的锚点片段
    return urlunparse(parsed._replace(fragment=''))

# 在get_links中替换链接处理逻辑
links = [normalize_url(urljoin(url, a['href'])) for a in soup.find_all('a', href=True)]

额外优化建议

  • 使用set存储visited,避免O(n)时间复杂度的查找操作。
  • 为requests.get添加timeout参数,防止请求长时间阻塞。
  • 可添加User-Agent请求头,避免被服务器识别为爬虫导致请求失败。

内容的提问来源于stack exchange,提问作者David Frick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 04:59:53