如何使用graph-tool将hSBM生成的分层二分图转换为仅含词与词组的圆形图?
Hey there! Let’s break down how to fix your visualization issue and get that circular word-topic graph you’re after. First, let’s clear up why your previous attempts didn’t work, then walk through the correct approach step by step.
Why Your Earlier Methods Failed
Let’s quickly diagnose the problems with your two tries:
- Anchor Node Approach: Replacing documents with a single anchor creates a star-shaped graph (all words pointing to the center), which doesn’t capture thematic relationships between words. You also lose the original hSBM topic block metadata when deleting document nodes.
- Manual Edge Addition & Projection: Your filter
vfilt=model.g.vp.kind.a == 0selects document nodes instead of word nodes, so you ended up projecting a document graph instead of a word graph. Plus, manually adding edges between neighbors is inefficient and doesn’t preserve meaningful co-occurrence weights.
The Correct Workflow
We’ll leverage graph-tool’s built-in tools to project the bipartite graph to words only, retain the hSBM topic assignments, and then visualize with a circular layout grouped by topics.
Step 1: Project the Bipartite Graph to a Word Co-Occurrence Graph
graph-tool has a project_graph function designed exactly for this—it creates a graph of word nodes where edges represent how often two words appear together in documents.
from graph_tool.all import * # Project the original bipartite graph (model.g) to word nodes (kind=1) # This creates a graph where edges have a "weight" property equal to co-occurrence count word_graph = project_graph(model.g, model.g.vp['kind'], target=1)
Step 2: Migrate hSBM Topic Block Metadata
Your trained hSBM already has topic assignments for word nodes. We need to copy this metadata to our new word graph. Assuming your original model stores word topics in a vertex property like membership:
# Extract word-to-topic mapping from the original bipartite graph original_word_nodes = [v for v in model.g.vertices() if model.g.vp['kind'][v] == 1] word_to_topic = { model.g.vp['name'][v]: model.g.vp['membership'][v] for v in original_word_nodes } # Add topic property to the new word graph word_graph.vp['topic'] = word_graph.new_vp("int") for v in word_graph.vertices(): word_graph.vp['topic'][v] = word_to_topic[word_graph.vp['name'][v]]
Step 3: Generate Circular Layout Grouped by Topic
To get a clean circular layout where words from the same topic are clustered together, sort the nodes by their topic before applying the circular layout:
# Sort vertices by their topic to group them in the layout sorted_vertices = sorted(word_graph.vertices(), key=lambda v: word_graph.vp['topic'][v]) # Create circular layout using the sorted node order pos = circular_layout(word_graph, order=sorted_vertices) # Assign unique colors to each topic for clarity num_topics = len(set(word_graph.vp['topic'].a)) topic_colors = random_colors(num_topics) vertex_fill_color = word_graph.new_vp("vector<float>") for v in word_graph.vertices(): vertex_fill_color[v] = topic_colors[word_graph.vp['topic'][v]] # Draw the graph graph_draw( word_graph, pos=pos, vertex_fill_color=vertex_fill_color, vertex_text=word_graph.vp['name'], vertex_font_size=9, edge_pen_width=word_graph.ep['weight'] / 15, # Scale edge width by co-occurrence count output_size=(1200, 1200), output="word_topic_circular.png" )
If You Don’t Need Co-Occurrence Edges
If you just want to visualize words grouped by topic (no edges between words), you can simplify further:
# Create a minimal graph with only word nodes simple_word_graph = Graph() simple_word_graph.add_vertex(len(original_word_nodes)) # Add name and topic properties simple_word_graph.vp['name'] = simple_word_graph.new_vp("string", vals=[model.g.vp['name'][v] for v in original_word_nodes]) simple_word_graph.vp['topic'] = simple_word_graph.new_vp("int", vals=[model.g.vp['membership'][v] for v in original_word_nodes]) # Generate sorted circular layout and draw sorted_v = sorted(simple_word_graph.vertices(), key=lambda v: simple_word_graph.vp['topic'][v]) pos = circular_layout(simple_word_graph, order=sorted_v) graph_draw( simple_word_graph, pos=pos, vertex_fill_color=vertex_fill_color, vertex_text=simple_word_graph.vp['name'], vertex_font_size=10, output="word_topic_simple_circle.png" )
Key Notes
- You don’t need to retrain the hSBM model—we’re just repurposing the topic assignments it already generated.
- Using
project_graphis far more efficient than manual edge creation, and it preserves meaningful co-occurrence weights for edges.
内容的提问来源于stack exchange,提问作者Jakob Rasmussen

