如何按特定分布将pandas DataFrame观测值分配至分组?
Got it, let's walk through how to solve this problem—assigning individuals to groups based on specific criteria, then building a network where within-group members connect at a given probability. I'll use pandas for grouping and networkx for network construction, since they're the go-to tools for this kind of work.
First, we'll take your pandas DataFrame and assign eligible individuals to groups (like schools for 6-10 year olds). Let's start with a concrete example.
First, import the necessary libraries:
import pandas as pd import numpy as np import networkx as nx from itertools import combinations # For efficient pair generation
Let's create a sample DataFrame to work with—100 kids aged 6-10, each with a unique ID:
np.random.seed(42) # Fix seed for reproducibility df = pd.DataFrame({ "id": range(1, 101), "age": np.random.randint(6, 11, size=100) })
Now, assign groups. For this example, we'll split the 6-10 year olds into 5 schools. If you only wanted to target a subset (say 6-8 year olds), you can use loc to filter:
# Assign all 6-10 year olds to 5 schools num_schools = 5 df["school_group"] = np.random.randint(1, num_schools + 1, size=len(df)) # If you only want to assign groups to 6-8 year olds: # df.loc[df["age"].between(6, 8), "school_group"] = np.random.randint(1, num_schools+1, size=len(df[df["age"].between(6,8)])) # For 9-10 year olds, you could assign a different type of group (like clubs): # df.loc[df["age"].between(9,10), "club_group"] = np.random.randint(1, 5, size=len(df[df["age"].between(9,10)]))
For more balanced groups (instead of random), you could use pd.qcut or split by ID buckets—adjust based on your needs!
Next, we'll create a network where members of the same group connect with probability p. The core idea is:
- Add all individuals as nodes to the network.
- For each group, generate all possible pairs of members.
- For each pair, flip a "coin" with probability
pto decide if an edge exists between them.
Here's a clean function to do this:
def build_group_network(df, group_column, connection_prob): # Initialize an empty undirected graph G = nx.Graph() # Add all individuals as nodes G.add_nodes_from(df["id"]) # Iterate over each group in the specified column for _, group_members in df.groupby(group_column): member_ids = group_members["id"].tolist() # Generate all unique pairs of members (no self-connections, no duplicates) for pair in combinations(member_ids, 2): # Randomly decide to add an edge based on the probability if np.random.random() < connection_prob: G.add_edge(*pair) return G
Let's use this function to build a school network with a 20% connection probability:
school_network = build_group_network(df, "school_group", connection_prob=0.2) # Check basic network stats print(f"Total nodes: {school_network.number_of_nodes()}") print(f"Total edges: {school_network.number_of_edges()}")
If you want to see how the groups and connections look, you can plot the network with color-coded groups:
import matplotlib.pyplot as plt # Map group IDs to colors for visualization color_map = {1: "red", 2: "blue", 3: "green", 4: "orange", 5: "purple"} node_colors = df.set_index("id")["school_group"].map(color_map) plt.figure(figsize=(10, 8)) nx.draw(school_network, node_color=node_colors, with_labels=True, node_size=500, font_size=8, alpha=0.7) plt.title("School Group Network (20% Within-Group Connection Probability)") plt.show()
- Complex Grouping: If you need to group by multiple criteria (e.g., age + gender), create a composite group column:
df["combined_group"] = df["gender"] + "_" + df["school_group"].astype(str) - Large Datasets: For very big DataFrames, nested loops might be slow. You can optimize by generating all possible edges for a group, then randomly sampling a fraction
pof them instead of checking each pair individually. - Reproducibility: Always set
np.random.seed()if you need consistent results across runs.
内容的提问来源于stack exchange,提问作者Wilco

