如何读取包含顶点分区的Pajek(.paj)格式文件?
读取含顶点分区的Pajek文件解决方案
问题根源
NetworkX的nx.read_pajek()方法仅读取Pajek文件中的顶点基本信息和边数据,不会解析顶点分区(*Partition块);你之前的自定义代码只针对*Vertices块做了处理,没有匹配分区相关的标记,导致分区数据未被捕获,最终返回空列表。
解决方案一:改进自定义解析代码
下面的代码可以完整读取顶点信息、边数据以及顶点分区:
import networkx as nx def read_pajek_with_partition(file_path): G = nx.DiGraph() # 无向图替换为nx.Graph() vertices = {} partitions = {} current_section = None current_partition_name = None with open(file_path, 'r') as f: for line in f: line = line.strip() if not line: continue # 识别区块标记 if line.startswith('*Vertices'): current_section = 'vertices' continue elif line.startswith('*Edges') or line.startswith('*Arcs'): current_section = 'edges' # 切换图的有向/无向属性 is_directed = line.startswith('*Arcs') if is_directed and not isinstance(G, nx.DiGraph): G = nx.DiGraph() elif not is_directed and isinstance(G, nx.DiGraph): G = nx.Graph() continue elif line.startswith('*Partition'): current_section = 'partition' # 提取分区名称(如*Partition gender -> 分区名gender) parts = line.split() current_partition_name = parts[1] if len(parts) > 1 else 'default_partition' partitions[current_partition_name] = {} continue # 处理各区块内容 if current_section == 'vertices': # 顶点行格式:ID "名称" [属性...] parts = line.split('"') vertex_id = parts[0].strip() vertex_name = parts[1].strip() if len(parts) > 1 else vertex_id extra_attrs = parts[2].split() if len(parts) > 2 else [] vertices[vertex_id] = {'name': vertex_name} if extra_attrs: vertices[vertex_id]['attributes'] = extra_attrs G.add_node(vertex_id, **vertices[vertex_id]) elif current_section == 'edges': # 边行格式:源ID 目标ID [权重] parts = line.split() source, target = parts[0], parts[1] weight = float(parts[2]) if len(parts) > 2 else 1.0 G.add_edge(source, target, weight=weight) elif current_section == 'partition': # 分区行格式:顶点ID 分区标签 parts = line.split() vertex_id, part_label = parts[0], parts[1] # 将分区信息写入节点属性 G.nodes[vertex_id][current_partition_name] = part_label partitions[current_partition_name][vertex_id] = part_label return G, partitions
使用示例
# 读取目标文件 graph, partition_dict = read_pajek_with_partition('path/file.paj') # 查看单个节点的完整属性(含分区) print(graph.nodes['1']) # 查看所有分区的映射关系 print(partition_dict)
解决方案二:使用第三方库
如果不想手动编写解析逻辑,可以使用专门处理Pajek文件的第三方库pajekreader(需先安装:pip install pajekreader),它原生支持解析分区等扩展信息:
from pajekreader import PajekReader reader = PajekReader('path/file.paj') graph = reader.get_network() partitions = reader.get_partitions() # 输出所有分区数据 print(partitions)
内容的提问来源于stack exchange,提问作者LdM
相关产品推荐
相关产品推荐

