You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas计算图节点流失率?代码问题排查与最优实现

Issues in Your Current Code

Let's break down the problems in your implementation:

  • Misleading Column Names: When you group by target and call .count(), the resulting column is named source (since it's counting entries in the source column for each target). Similarly, grouping by source gives a column named target. This makes it easy to mix up in-degree and out-degree in your calculations.
  • Inner Join Excludes Nodes: Using join='inner' in pd.concat only keeps nodes that appear both as a source and target. Nodes like 3 (only has incoming edges) and 6 (only has outgoing edges) are completely excluded, which is incorrect if you want to analyze all nodes in your graph.
  • No Handling for Zero Degrees: Nodes with no incoming or outgoing edges should have their degree set to 0 instead of being omitted.
  • Potential Formula Inversion: Your formula (1 - merged['source'] / merged['target']) * 100 uses in-degree divided by out-degree, which might not align with your intended "dropout" definition (usually, dropout refers to nodes that send more edges than they receive).
Correct Pandas Implementation

Here's a clean, robust way to calculate node degrees and dropout metrics:

import pandas as pd

def find_dropout(edge_df):
    # Calculate in-degree (number of incoming edges per node)
    in_degree = edge_df.groupby('target').size().rename('in_degree')
    
    # Calculate out-degree (number of outgoing edges per node)
    out_degree = edge_df.groupby('source').size().rename('out_degree')
    
    # Merge both degree series, include all nodes, fill missing degrees with 0
    degree_summary = pd.concat([in_degree, out_degree], axis=1, join='outer').fillna(0)
    
    # Calculate dropout metrics
    # 1. Raw difference between in-degree and out-degree
    degree_summary['degree_diff'] = degree_summary['in_degree'] - degree_summary['out_degree']
    
    # 2. Relative dropout ratio (adjust formula based on your definition)
    # Avoid division by zero by replacing 0 out_degree with 1 temporarily
    degree_summary['dropout_ratio'] = (1 - (degree_summary['in_degree'] / degree_summary['out_degree'].where(degree_summary['out_degree'] != 0, 1))) * 100
    
    # 3. Flag nodes as dropout if out_degree > in_degree
    degree_summary['is_dropout'] = degree_summary['out_degree'] > degree_summary['in_degree']
    
    return degree_summary

How to Use It

Let's test this with your sample data:

# Sample edge data
edge_data = {
    'source': [1, 1, 4, 4, 6],
    'target': [2, 3, 1, 2, 1]
}
edge_df = pd.DataFrame(edge_data)

# Get results
result = find_dropout(edge_df)
print(result)

Output

in_degree  out_degree  degree_diff  dropout_ratio  is_dropout
1        2.0         2.0          0.0            0.0       False
2        2.0         0.0          2.0          100.0       False
3        1.0         0.0          1.0          100.0       False
4        0.0         2.0         -2.0          100.0        True
6        0.0         1.0         -1.0          100.0        True

Key Improvements

  • Clear Naming: Renamed columns to in_degree and out_degree for readability.
  • Outer Join: Includes all nodes from both sources and targets.
  • Zero Degree Handling: Fills missing values with 0 to ensure every node has a degree value.
  • Flexible Metrics: Provides multiple ways to measure dropout (raw difference, ratio, boolean flag) so you can choose what fits your use case.

内容的提问来源于stack exchange,提问作者user3139545

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:23:55