You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于引用网络的Pandas DataFrame New_var字段生成需求

Pandas 数据处理:构建New_var字段实现引用节点Fields值映射

需求说明

  • 数据集包含字段:docdb_id(节点ID)、cited_patents(当前节点引用的节点列表)、dist_cited_patents(对应引用节点的距离列表)、Fields(节点的字段信息)
  • 节点距离规则:docdb_id的距离为min(dist_cited_patents) + 1;Fields不为空的节点,距离视为0
  • 需新增New_var字段:提取dist_cited_patents中最小距离对应的所有cited_patents的Fields值,合并后作为该字段内容

原始数据集

docdb_id Cited_patents Dist_cited_patents   Fields. 
1              [7,3]            [1,1]            ""     
2              [1,5]            [2,1]           ['Math']    
3              [1,2,6]          [2,0,2]          ""     
4              [7]              [1]             ['1. Natural Sciences' '3. Medical and Health Sciences']
5              [1,2]            [2,0]             ""      
6              [5,8]            [1,1]             ""       
7              [4,8]            [0,1]            ""      
8              [4]              [0]              ""      

期望结果

docdb_id Cited_patents Dist_cited_patents   Fields     New_var
1              [7,3]            [1,1]        ""    [Natural Sciences, Medical and Health Sciences, Math]
2              [1,5]            [2,1]        ['Math']       Math
3              [1,2,6]          [2,0,2]      ""         Math
4              [7]              [1]         ['1. Natural Sciences' '3. Medical and Health Sciences']    Natural Sciences, Medical and Health Sciences,
5              [1,2]            [2,0]         ""        Math
6              [5,8]            [1,1]         ""    [Natural Sciences, Medical and Health Sciences,, Math]  
7              [4,8]            [0,1]         ""      Science 
8              [4]              [0]           ""      Science

数据集初始化代码

import pandas as pd

# 初始化数据列表
data = [
    [1, [7,3], [1,1], ""], 
    [2, [1,5], [2,1], "Math"], 
    [3, [1,2,6], [2,0,2], ""],
    [4, [7], [1], "Science"],
    [5, [1,2], [2,0], ""],
    [6, [5,8], [1,1], ""],
    [7, [4,8], [0,1], ""],
    [8, [4], [0], ""]
]
  
# 创建DataFrame
df = pd.DataFrame(data, columns=['docdb', 'cited_patents','dist_cited_patents','Fields'])

解决方案代码

步骤1:修正列名并创建节点-字段映射字典

统一列名并构建快速查询引用节点字段的映射:

# 修正列名
df.rename(columns={'docdb': 'docdb_id'}, inplace=True)

# 创建docdb_id到Fields的映射字典,空字符串保留原格式
field_map = df.set_index('docdb_id')['Fields'].to_dict()

步骤2:定义生成New_var的函数

遍历每行数据,筛选最小距离对应的引用节点并提取字段:

def generate_new_var(row):
    dists = row['dist_cited_patents']
    cited_nodes = row['cited_patents']
    
    # 找到当前行的最小距离
    min_dist = min(dists)
    # 筛选所有距离等于最小值的引用节点
    target_nodes = [node for dist, node in zip(dists, cited_nodes) if dist == min_dist]
    # 提取对应节点的Fields值
    target_fields = [field_map[node] for node in target_nodes]
    
    # 匹配期望输出格式:单个值直接返回,多值用逗号分隔或方括号包裹
    if len(target_fields) == 1:
        return target_fields[0]
    else:
        return f"[{', '.join(target_fields)}]"

步骤3:应用函数生成New_var字段

df['New_var'] = df.apply(generate_new_var, axis=1)

若需要更贴合期望结果的格式细节(比如保留空值占位),可微调target_fields的处理逻辑,例如保留空字符串的占位:

target_fields = [field_map[node] if field_map[node] else '' for node in target_nodes]

内容的提问来源于stack exchange,提问作者Lusian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 11:35:23