You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于营收和参数数量优先级对Website列去重?

按规则对Website列去重的实现方案

可以用Python的pandas库实现这个需求,核心思路是先添加辅助排序字段,再按规则排序后分组取首行,具体步骤如下:

步骤说明

  1. 构造辅助字段:计算每行Facebook和Twitter列的和,命名为social_count,用来量化社交参数的数量
  2. 按规则排序:先按Revenue降序排列,再按social_count降序排列,确保同一Website的行中,优先级高的排在前面
  3. 分组去重:按Website分组,每组只保留第一行数据

完整代码示例

import pandas as pd

# 构造示例数据集
data = {
    'Website': ['xyz.com', 'abc.com', 'xyz.com', 'def.com', 'rte.com', 'abc.com'],
    'Revenue': [345, 634, 345, 732, 298, 553],
    'Company_Age': [45, 42, 55, 63, 29, 20],
    'Facebook': [0, 0, 1, 1, 0, 0],
    'Twitter': [0, 0, 0, 0, 1, 1],
    'CEO': [pd.NA]*6
}
df = pd.DataFrame(data)

# 计算社交参数数量
df['social_count'] = df['Facebook'] + df['Twitter']

# 按规则排序:先营收降序,再社交参数数量降序
sorted_df = df.sort_values(by=['Revenue', 'social_count'], ascending=[False, False])

# 按Website分组,取每组第一行
result_df = sorted_df.groupby('Website', as_index=False).first()

# 删除辅助列(可选操作)
result_df = result_df.drop('social_count', axis=1)

print(result_df)

输出结果

运行后得到符合规则的去重数据:

Website  Revenue  Company_Age  Facebook  Twitter   CEO
0  abc.com      634           42         0        0  <NA>
1  def.com      732           63         1        0  <NA>
2  rte.com      298           29         0        1  <NA>
3  xyz.com      345           55         1        0  <NA>

内容的提问来源于stack exchange,提问作者Patrik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 01:52:46