基于引用网络的Pandas DataFrame New_var字段生成需求
Pandas 数据处理:构建New_var字段实现引用节点Fields值映射
需求说明
- 数据集包含字段:
docdb_id(节点ID)、cited_patents(当前节点引用的节点列表)、dist_cited_patents(对应引用节点的距离列表)、Fields(节点的字段信息) - 节点距离规则:
docdb_id的距离为min(dist_cited_patents) + 1;Fields不为空的节点,距离视为0 - 需新增
New_var字段:提取dist_cited_patents中最小距离对应的所有cited_patents的Fields值,合并后作为该字段内容
原始数据集
docdb_id Cited_patents Dist_cited_patents Fields. 1 [7,3] [1,1] "" 2 [1,5] [2,1] ['Math'] 3 [1,2,6] [2,0,2] "" 4 [7] [1] ['1. Natural Sciences' '3. Medical and Health Sciences'] 5 [1,2] [2,0] "" 6 [5,8] [1,1] "" 7 [4,8] [0,1] "" 8 [4] [0] ""
期望结果
docdb_id Cited_patents Dist_cited_patents Fields New_var 1 [7,3] [1,1] "" [Natural Sciences, Medical and Health Sciences, Math] 2 [1,5] [2,1] ['Math'] Math 3 [1,2,6] [2,0,2] "" Math 4 [7] [1] ['1. Natural Sciences' '3. Medical and Health Sciences'] Natural Sciences, Medical and Health Sciences, 5 [1,2] [2,0] "" Math 6 [5,8] [1,1] "" [Natural Sciences, Medical and Health Sciences,, Math] 7 [4,8] [0,1] "" Science 8 [4] [0] "" Science
数据集初始化代码
import pandas as pd # 初始化数据列表 data = [ [1, [7,3], [1,1], ""], [2, [1,5], [2,1], "Math"], [3, [1,2,6], [2,0,2], ""], [4, [7], [1], "Science"], [5, [1,2], [2,0], ""], [6, [5,8], [1,1], ""], [7, [4,8], [0,1], ""], [8, [4], [0], ""] ] # 创建DataFrame df = pd.DataFrame(data, columns=['docdb', 'cited_patents','dist_cited_patents','Fields'])
解决方案代码
步骤1:修正列名并创建节点-字段映射字典
统一列名并构建快速查询引用节点字段的映射:
# 修正列名 df.rename(columns={'docdb': 'docdb_id'}, inplace=True) # 创建docdb_id到Fields的映射字典,空字符串保留原格式 field_map = df.set_index('docdb_id')['Fields'].to_dict()
步骤2:定义生成New_var的函数
遍历每行数据,筛选最小距离对应的引用节点并提取字段:
def generate_new_var(row): dists = row['dist_cited_patents'] cited_nodes = row['cited_patents'] # 找到当前行的最小距离 min_dist = min(dists) # 筛选所有距离等于最小值的引用节点 target_nodes = [node for dist, node in zip(dists, cited_nodes) if dist == min_dist] # 提取对应节点的Fields值 target_fields = [field_map[node] for node in target_nodes] # 匹配期望输出格式:单个值直接返回,多值用逗号分隔或方括号包裹 if len(target_fields) == 1: return target_fields[0] else: return f"[{', '.join(target_fields)}]"
步骤3:应用函数生成New_var字段
df['New_var'] = df.apply(generate_new_var, axis=1)
若需要更贴合期望结果的格式细节(比如保留空值占位),可微调target_fields的处理逻辑,例如保留空字符串的占位:
target_fields = [field_map[node] if field_map[node] else '' for node in target_nodes]
内容的提问来源于stack exchange,提问作者Lusian
相关产品推荐
相关产品推荐

