You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用NetworkX提升Python中读取TXT文件的Graph()函数效率

优化NetworkX读取15万行关联数据的效率

Great question—handling 150k lines efficiently with NetworkX definitely needs some targeted optimizations, since naive line-by-line processing can get slow thanks to repeated function calls and string parsing overhead. Here are several actionable tweaks to speed things up:

1. 预编译正则表达式,减少字符串解析耗时

Your line format follows a predictable pattern: #X - string #Y #Z.... The key is to extract the starting node ID and all linked IDs efficiently. Precompiling a regular expression avoids re-compiling the pattern for every single line, which saves a ton of time when processing hundreds of thousands of lines.

示例代码:

import re
import networkx as nx

# 预编译正则:匹配所有带#的数字ID
id_pattern = re.compile(r'#(\d+)')

G = nx.Graph()
edge_list = []

# 加大文件缓冲区,减少IO交互次数
with open('your_data.txt', 'r', buffering=1024*1024) as f:
    for line in f:
        # 提取当前行所有数字ID
        all_ids = list(map(int, id_pattern.findall(line)))
        if len(all_ids) < 2:
            continue  # 跳过无关联节点的行
        source_node = all_ids[0]
        # 生成当前节点与所有关联节点的边
        edge_list.extend([(source_node, target) for target in all_ids[1:]])

# 批量添加所有边,比逐行add_edge高效得多
G.add_edges_from(edge_list)

2. 用批量操作替代逐行操作

NetworkX's add_edges_from is way more efficient than calling add_edge hundreds of thousands of times. This is because batch operations minimize the overhead of switching between Python and NetworkX's underlying C extensions. Collecting all edges first and adding them in one go can cut processing time by a factor of 5-10.

3. 优化文件读取逻辑

Setting a large buffering value (like 1MB in the example) reduces the number of times your code interacts with the filesystem—critical for large files. If you don't need the actual string content between IDs, you could also split lines directly, but regex is more reliable for messy or variable formatting.

4. 可选:先用轻量结构预处理(如果不需要NetworkX全功能)

If your only needs are basic graph operations (like querying neighbors) and you don't need NetworkX's advanced algorithms, using a Python dict to build an adjacency table first can be even faster. You can convert it to a NetworkX graph later if needed:

import re

id_pattern = re.compile(r'#(\d+)')
adjacency_table = {}

with open('your_data.txt', 'r', buffering=1024*1024) as f:
    for line in f:
        all_ids = list(map(int, id_pattern.findall(line)))
        if not all_ids:
            continue
        source = all_ids[0]
        targets = all_ids[1:]
        # 无向图需要双向添加关联
        if source not in adjacency_table:
            adjacency_table[source] = set()
        adjacency_table[source].update(targets)
        for target in targets:
            if target not in adjacency_table:
                adjacency_table[target] = set()
            adjacency_table[target].add(source)

# 转成NetworkX图(如果需要)
G = nx.Graph()
for node, neighbors in adjacency_table.items():
    G.add_edges_from([(node, neighbor) for neighbor in neighbors])

5. 环境层面:用PyPy替代CPython

If your code can run on PyPy, its JIT compiler will drastically speed up loops and string processing. For 150k lines, you might see a 2-5x speedup compared to standard CPython.


内容的提问来源于stack exchange,提问作者Yafim Simanovsky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:01:33