如何设置参考行,使DataFrame中其他行对应值除以参考行的值?
问题描述
我有一个包含不同细胞及其mRNA表达数据的大型DataFrame,结构如下:
| A | Gene 1 | Gene 2 | Gene 3 | Gene 4 | … | |
|---|---|---|---|---|---|---|
| Cell 1 | 9 | 12 | 24 | 42 | 30 | |
| Cell 2 | 3 | 6 | 12 | 21 | 15 | |
| Cell 3 | 6 | 42 | 48 | 84 | 45 | |
| … |
需要将指定行(比如示例中的Cell 2)设为参考行,让其他行的各列数值除以参考行对应列的数值,得到如下标准化结果:
| A | Gene 1 | Gene 2 | Gene 3 | Gene 4 | … | |
|---|---|---|---|---|---|---|
| Cell 1 | 3 | 2 | 2 | 2 | 2 | |
| Cell 2 | 1 | 1 | 1 | 1 | 1 | |
| Cell 3 | 2 | 7 | 4 | 4 | 3 | |
| … |
解决方案(基于Pandas)
假设你的DataFrame已加载完成,且细胞名称是索引列,可以通过以下简单步骤实现:
提取参考行数值
先获取指定参考行的所有列数值,以Cell 2为例:ref_values = df.loc["Cell 2"]执行标准化计算
Pandas会自动按列对齐,直接让整个DataFrame除以参考行的Series即可:normalized_df = df / ref_values(可选)转换为整数格式
如果需要和示例一样保留整数结果,可转换数据类型:normalized_df = normalized_df.astype(int)
完整示例代码
import pandas as pd # 构造示例DataFrame data = { "A": [9, 3, 6], "Gene 1": [12, 6, 42], "Gene 2": [24, 12, 48], "Gene 3": [42, 21, 84], "Gene 4": [30, 15, 45] } df = pd.DataFrame(data, index=["Cell 1", "Cell 2", "Cell 3"]) # 指定参考行 ref_row = "Cell 2" ref_values = df.loc[ref_row] # 标准化计算 normalized_df = df / ref_values # 转换为整数(按需选择) normalized_df = normalized_df.astype(int) print(normalized_df)
运行后输出:
A Gene 1 Gene 2 Gene 3 Gene 4 Cell 1 3 2 2 2 2 Cell 2 1 1 1 1 1 Cell 3 2 7 4 4 3
注意事项
- 确保参考行中没有
0值,否则会触发除以0的错误;若存在0,需提前处理(如跳过对应列或填充默认值)。 - 如果DataFrame使用数字索引,直接用索引值提取参考行即可,比如
ref_values = df.loc[1](假设Cell 2对应索引1)。
内容的提问来源于stack exchange,提问作者NPC
相关产品推荐
相关产品推荐

