You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas中为不同DataFrame的同名数据分配相同ID的实现

为关联数据集生成唯一ID并使用Featuretools进行深度特征合成

1. 生成并匹配唯一ID

假设你的两个数据集分别为unique_players(唯一名称数据集)和season_stats(含重复同名记录的数据集),可通过以下步骤完成ID的生成与匹配:

步骤1:为唯一名称数据集添加唯一ID

如果unique_players尚未包含player_id列,直接生成自增ID即可:

import pandas as pd

# 构造示例唯一名称数据集
unique_players = pd.DataFrame({
    "Name": ["John Dosh", "Michael Deesh", "Julia Roberts"]
})

# 生成唯一player_id
unique_players["player_id"] = range(1, len(unique_players) + 1)

步骤2:为重复记录数据集匹配对应ID

通过merge操作,基于Name字段将season_stats与unique_players关联,同步对应player_id:

# 构造示例含重复记录的数据集
season_stats = pd.DataFrame({
    "Name": ["John Dosh", "John Dosh", "Michael Deesh", "Michael Deesh", "Michael Deesh", "Julia Roberts", "Julia Roberts"]
})

# 匹配对应player_id
season_stats = pd.merge(season_stats, unique_players[["Name", "player_id"]], on="Name", how="left")

# 为season_stats添加唯一索引列(Featuretools要求实体必须有唯一索引)
season_stats["season_stats_id"] = range(1, len(season_stats) + 1)

处理后的数据将与你提供的示例结构完全一致。

2. 使用Featuretools构建实体集并创建关联关系

完成ID匹配后,即可按以下代码构建实体集、创建实体关联,为后续深度特征合成做准备:

import featuretools as ft

# 创建实体集
entity_set = ft.EntitySet("basketball_players")

# 添加唯一球员数据集实体
entity_set.add_dataframe(
    dataframe_name="players_set",
    dataframe=unique_players,
    index="player_id"  # 推荐使用唯一ID作为实体索引,比名称更可靠
)

# 添加赛季统计数据集实体
entity_set.add_dataframe(
    dataframe_name="season_stats",
    dataframe=season_stats,
    index="season_stats_id"
)

# 建立两个实体间的关联关系
entity_set.add_relationship(
    "players_set",  # 父实体名称
    "player_id",    # 父实体关联字段
    "season_stats", # 子实体名称
    "player_id"     # 子实体关联字段
)

注:原示例中players_set的index使用了name,但实际更推荐用唯一的player_id作为索引,可避免名称拼写错误、重复等潜在问题。

内容的提问来源于stack exchange,提问作者idunskyi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 23:20:28