You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并DataFrame行:实现邮箱唯一关联所有标记系统

合并重复邮箱并关联所有标记系统的解决方案

嘿,看起来你需要把重复邮箱对应的标记系统合并到一行里对吧?我给你几个实用的方案,根据你日常使用的工具来选就好~

方案1:用Python代码处理(适合编程场景)

假设你的输出内容是每行以 邮箱 | 标记系统 的格式呈现,比如:

user1@example.com | SystemA
user1@example.com | SystemB
user2@example.com | SystemA

你可以用字典来聚合每个邮箱对应的所有标记系统,代码示例如下:

# 这里可以替换成你实际获取输出行的方式,比如从文件读取:
# with open("your_output.txt", "r") as f:
#     output_lines = f.readlines()

output_lines = [
    "user1@example.com | SystemA",
    "user1@example.com | SystemB",
    "user2@example.com | SystemA"
]

email_systems = {}
for line in output_lines:
    # 去除首尾空格后分割行内容
    email, system = line.strip().split(" | ")
    # 聚合标记系统
    if email not in email_systems:
        email_systems[email] = []
    email_systems[email].append(system)

# 生成你期望的输出格式
for email, systems in email_systems.items():
    print(f"{email} | {', '.join(systems)}")

运行后就能得到:

user1@example.com | SystemA, SystemB
user2@example.com | SystemA

如果你的原输出分隔符不是 |,只需要修改 split 里的参数就行。

方案2:用Shell脚本(awk)处理(适合文本文件快速处理)

如果你的输出已经保存在文本文件里(比如 output.txt),可以直接用awk命令快速完成合并:

awk -F' | ' '{arr[$1] = arr[$1] ? arr[$1] ", " $3 : $3} END{for (e in arr) print e " | " arr[e]}' output.txt

命令解释:

  • -F' | ':指定行内容的分隔符为 |(空格+竖线+空格)
  • arr[$1] = arr[$1] ? arr[$1] ", " $3 : $3:用数组存储邮箱对应的标记系统,已存在则追加新系统,不存在则初始化
  • END 块:遍历数组,输出每个邮箱和对应的所有标记系统

内容的提问来源于stack exchange,提问作者BeMuRR187

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:18:16