You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取包含顶点分区的Pajek(.paj)格式文件?

读取含顶点分区的Pajek文件解决方案

问题根源

NetworkX的nx.read_pajek()方法仅读取Pajek文件中的顶点基本信息和边数据,不会解析顶点分区(*Partition块);你之前的自定义代码只针对*Vertices块做了处理,没有匹配分区相关的标记,导致分区数据未被捕获,最终返回空列表。

解决方案一:改进自定义解析代码

下面的代码可以完整读取顶点信息、边数据以及顶点分区:

import networkx as nx

def read_pajek_with_partition(file_path):
    G = nx.DiGraph()  # 无向图替换为nx.Graph()
    vertices = {}
    partitions = {}
    current_section = None
    current_partition_name = None

    with open(file_path, 'r') as f:
        for line in f:
            line = line.strip()
            if not line:
                continue
            
            # 识别区块标记
            if line.startswith('*Vertices'):
                current_section = 'vertices'
                continue
            elif line.startswith('*Edges') or line.startswith('*Arcs'):
                current_section = 'edges'
                # 切换图的有向/无向属性
                is_directed = line.startswith('*Arcs')
                if is_directed and not isinstance(G, nx.DiGraph):
                    G = nx.DiGraph()
                elif not is_directed and isinstance(G, nx.DiGraph):
                    G = nx.Graph()
                continue
            elif line.startswith('*Partition'):
                current_section = 'partition'
                # 提取分区名称(如*Partition gender -> 分区名gender)
                parts = line.split()
                current_partition_name = parts[1] if len(parts) > 1 else 'default_partition'
                partitions[current_partition_name] = {}
                continue
            
            # 处理各区块内容
            if current_section == 'vertices':
                # 顶点行格式:ID "名称" [属性...]
                parts = line.split('"')
                vertex_id = parts[0].strip()
                vertex_name = parts[1].strip() if len(parts) > 1 else vertex_id
                extra_attrs = parts[2].split() if len(parts) > 2 else []
                vertices[vertex_id] = {'name': vertex_name}
                if extra_attrs:
                    vertices[vertex_id]['attributes'] = extra_attrs
                G.add_node(vertex_id, **vertices[vertex_id])
            
            elif current_section == 'edges':
                # 边行格式:源ID 目标ID [权重]
                parts = line.split()
                source, target = parts[0], parts[1]
                weight = float(parts[2]) if len(parts) > 2 else 1.0
                G.add_edge(source, target, weight=weight)
            
            elif current_section == 'partition':
                # 分区行格式:顶点ID 分区标签
                parts = line.split()
                vertex_id, part_label = parts[0], parts[1]
                # 将分区信息写入节点属性
                G.nodes[vertex_id][current_partition_name] = part_label
                partitions[current_partition_name][vertex_id] = part_label
    
    return G, partitions

使用示例

# 读取目标文件
graph, partition_dict = read_pajek_with_partition('path/file.paj')

# 查看单个节点的完整属性(含分区)
print(graph.nodes['1'])

# 查看所有分区的映射关系
print(partition_dict)

解决方案二:使用第三方库

如果不想手动编写解析逻辑,可以使用专门处理Pajek文件的第三方库pajekreader(需先安装:pip install pajekreader),它原生支持解析分区等扩展信息:

from pajekreader import PajekReader

reader = PajekReader('path/file.paj')
graph = reader.get_network()
partitions = reader.get_partitions()

# 输出所有分区数据
print(partitions)

内容的提问来源于stack exchange,提问作者LdM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 23:05:27