You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按特定分布将pandas DataFrame观测值分配至分组?

Got it, let's walk through how to solve this problem—assigning individuals to groups based on specific criteria, then building a network where within-group members connect at a given probability. I'll use pandas for grouping and networkx for network construction, since they're the go-to tools for this kind of work.

Step 1: Assign Individuals to Target Groups

First, we'll take your pandas DataFrame and assign eligible individuals to groups (like schools for 6-10 year olds). Let's start with a concrete example.

First, import the necessary libraries:

import pandas as pd
import numpy as np
import networkx as nx
from itertools import combinations  # For efficient pair generation

Let's create a sample DataFrame to work with—100 kids aged 6-10, each with a unique ID:

np.random.seed(42)  # Fix seed for reproducibility
df = pd.DataFrame({
    "id": range(1, 101),
    "age": np.random.randint(6, 11, size=100)
})

Now, assign groups. For this example, we'll split the 6-10 year olds into 5 schools. If you only wanted to target a subset (say 6-8 year olds), you can use loc to filter:

# Assign all 6-10 year olds to 5 schools
num_schools = 5
df["school_group"] = np.random.randint(1, num_schools + 1, size=len(df))

# If you only want to assign groups to 6-8 year olds:
# df.loc[df["age"].between(6, 8), "school_group"] = np.random.randint(1, num_schools+1, size=len(df[df["age"].between(6,8)]))
# For 9-10 year olds, you could assign a different type of group (like clubs):
# df.loc[df["age"].between(9,10), "club_group"] = np.random.randint(1, 5, size=len(df[df["age"].between(9,10)]))

For more balanced groups (instead of random), you could use pd.qcut or split by ID buckets—adjust based on your needs!

Step 2: Build the Network with Within-Group Connection Probabilities

Next, we'll create a network where members of the same group connect with probability p. The core idea is:

  1. Add all individuals as nodes to the network.
  2. For each group, generate all possible pairs of members.
  3. For each pair, flip a "coin" with probability p to decide if an edge exists between them.

Here's a clean function to do this:

def build_group_network(df, group_column, connection_prob):
    # Initialize an empty undirected graph
    G = nx.Graph()
    # Add all individuals as nodes
    G.add_nodes_from(df["id"])
    
    # Iterate over each group in the specified column
    for _, group_members in df.groupby(group_column):
        member_ids = group_members["id"].tolist()
        # Generate all unique pairs of members (no self-connections, no duplicates)
        for pair in combinations(member_ids, 2):
            # Randomly decide to add an edge based on the probability
            if np.random.random() < connection_prob:
                G.add_edge(*pair)
    return G

Let's use this function to build a school network with a 20% connection probability:

school_network = build_group_network(df, "school_group", connection_prob=0.2)

# Check basic network stats
print(f"Total nodes: {school_network.number_of_nodes()}")
print(f"Total edges: {school_network.number_of_edges()}")
Step 3: Optional - Visualize the Network

If you want to see how the groups and connections look, you can plot the network with color-coded groups:

import matplotlib.pyplot as plt

# Map group IDs to colors for visualization
color_map = {1: "red", 2: "blue", 3: "green", 4: "orange", 5: "purple"}
node_colors = df.set_index("id")["school_group"].map(color_map)

plt.figure(figsize=(10, 8))
nx.draw(school_network, 
        node_color=node_colors, 
        with_labels=True, 
        node_size=500, 
        font_size=8,
        alpha=0.7)
plt.title("School Group Network (20% Within-Group Connection Probability)")
plt.show()
Quick Tips for Edge Cases
  • Complex Grouping: If you need to group by multiple criteria (e.g., age + gender), create a composite group column: df["combined_group"] = df["gender"] + "_" + df["school_group"].astype(str)
  • Large Datasets: For very big DataFrames, nested loops might be slow. You can optimize by generating all possible edges for a group, then randomly sampling a fraction p of them instead of checking each pair individually.
  • Reproducibility: Always set np.random.seed() if you need consistent results across runs.

内容的提问来源于stack exchange,提问作者Wilco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:21:17