如何用NetworkX为节点创建边?并计算输入文件中蛋白质的度
Hey there! Let's tackle your two NetworkX questions step by step—first, we'll fix the small issue in your existing node creation code, then add edges from your input file, and finally calculate each protein's degree.
Fixing Node Creation & Adding Edges
First off, your original code has a tiny bug: you're repeatedly adding a node named 1 instead of using the actual protein identifiers from your data. Let's fix that, plus properly parse your input file into edge pairs (since your file content A B a b c d a c c b looks like sequential node pairs: A-B, a-b, c-d, a-c, c-b).
Here's the revised code with explanations:
import pandas as pd import networkx as nx # Read the input file and split all elements into a list with open('protein.txt', 'r') as f: all_elements = f.read().strip().split() # Convert the list into edge pairs (take every 2 elements as a pair) edge_list = list(zip(all_elements[::2], all_elements[1::2])) # Initialize the graph and add edges G = nx.Graph() G.add_edges_from(edge_list) # NetworkX automatically adds nodes for any new entries in edges!
A quick note: You don't need to manually add nodes first—add_edges_from() will automatically create any nodes that aren't already in the graph. If you wanted to explicitly add nodes (e.g., if some proteins have no edges), you could still use your distinct list like this:
distinct_proteins = list(set(all_elements)) G.add_nodes_from(distinct_proteins)
Calculating Protein Degrees
Once your graph is built, getting each protein's degree is super straightforward with NetworkX's built-in nx.degree() function. You can convert the results to a dictionary or DataFrame for easy reading:
# Calculate degrees and convert to a dictionary protein_degrees = dict(nx.degree(G)) # Print the results in a readable format print("Protein Degree Breakdown:") for protein, degree in protein_degrees.items(): print(f"- {protein}: {degree}") # Optional: Convert to a DataFrame for cleaner tabular output degree_df = pd.DataFrame.from_dict(protein_degrees, orient='index', columns=['Degree']) print("\nDegree DataFrame:") print(degree_df)
For your input data, the degrees will be:
- A: 1
- B: 2
- a: 2
- b: 2
- c: 3
- d: 1
That's all there is to it! The main fixes were using actual protein names as nodes and correctly parsing your input into edge pairs. NetworkX handles the rest smoothly.
内容的提问来源于stack exchange,提问作者AST

