You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python DataFrame技术问询:跨列元素查找与反应关系构建

我来帮你一步步解决这两个Pandas实用场景问题~


问题1:找出DataFrame某一列中同时出现在另一列的元素

先拿一个示例DataFrame来演示:

import pandas as pd

df = pd.DataFrame({
    'col1': ['A', 'B', 'C', 'D', 'E'],
    'col2': ['C', 'D', 'F', 'G', 'A']
})

要筛选col1里同时存在于col2的元素,用isin()方法就能轻松实现:

# 提取col1中在col2里的所有元素(含重复)
common_elements = df[df['col1'].isin(df['col2'])]['col1']

# 如果需要去重后的结果,加个.unique()或者.drop_duplicates()
unique_common = common_elements.unique().tolist()
print(unique_common)  # 输出: ['A', 'C', 'D']

如果你需要的是这些元素对应的整行数据,直接去掉后面的['col1']即可:

common_rows = df[df['col1'].isin(df['col2'])]

要是想找只在col1不在col2的元素,用~取反就行:

only_in_col1 = df[~df['col1'].isin(df['col2'])]['col1']

问题2:构建化学反应间的关联DataFrame

首先我先假设你的原始反应DataFrame结构是这样的(如果和你的实际结构有出入,调整字段名即可):

# 模拟你的反应数据结构
reaction_df = pd.DataFrame({
    'reaction_id': ['R1', 'R2'],
    'source': [['A', 'B'], ['C', 'D']],  # 反应物列表
    'target': [['C'], ['E']]             # 产物列表
})

我们的核心目标是找到反应X的产物是反应Y的反应物的关联,具体步骤如下:

步骤1:展开反应物/产物列表

因为每个反应可能有多个反应物或产物,先把列表拆成单独行,让每个化合物对应一个反应ID:

# 展开产物列:每个产物对应生成它的反应
target_expanded = reaction_df.explode('target').rename(columns={'target': 'compound', 'reaction_id': 'source_reaction'})

# 展开反应物列:每个反应物对应消耗它的反应
source_expanded = reaction_df.explode('source').rename(columns={'source': 'compound', 'reaction_id': 'target_reaction'})

步骤2:匹配反应间的关联

通过compound字段关联两个表,就能找到反应间的关联关系:

# 关联表,找出反应对
reaction_links = pd.merge(
    target_expanded[['source_reaction', 'compound']],
    source_expanded[['target_reaction', 'compound']],
    on='compound'
)

# 过滤掉反应自关联的情况(如果有的话)
reaction_links = reaction_links[reaction_links['source_reaction'] != reaction_links['target_reaction']]

# 去掉重复的关联对
reaction_links = reaction_links.drop_duplicates(subset=['source_reaction', 'target_reaction'])

步骤3:查看最终结果

运行后得到的reaction_links就是你需要的反应-反应关联DataFrame:

print(reaction_links)
# 输出:
#   source_reaction compound target_reaction
# 0              R1        C              R2

如果你的原始DataFrame里,source和target是逗号分隔的字符串(比如"A,B"),先转成列表再处理:

reaction_df['source'] = reaction_df['source'].str.split(',')
reaction_df['target'] = reaction_df['target'].str.split(',')

这样就能完美得到你要的反应关联关系啦~


内容的提问来源于stack exchange,提问作者Microdot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:41:42