You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何对不同DataFrame的多份样本取均值实现基因遗传预测

报错原因

  1. 核心错在df1[x] = df1.sample(n=num)这行逻辑:df.sample(n=num)返回的是包含num行、和原DataFrame列数一致的二维表,而你试图把这个二维表赋值给df1的单个新列,单条列要求是长度和原df行数一致的一维数据,形状完全不匹配,因此触发ValueError。
  2. 采样逻辑写错:你写的df2[x] = df1.sample(n=num)本来要采样第二个亲本的基因,写成了采样第一个亲本的数据,逻辑错误。

修正后的代码

你要实现的「从双方各采样50%、多轮采样规避异常值」的需求,可以参考以下实现:

import pandas as pd
import random

# 你原来的亲本数据构造代码
M = ("England & Northwestern Europe," * 53) + ("Ireland," * 21) + ("Scotland," * 21) + ("Wales," * 5)
M = M.split(',')
M.pop()

E = ("European Jewish," * 53) + ("Southern Italy," * 31) + ("Levant," * 8) + ("Northern Africa," * 3) + ("Aegean Islands," * 2) + ("Cyprus," * 2) + ("Arabian Peninsula," * 1)
E = E.split(',')
E.pop()

def simulate_baby_ancestry(parent1, parent2, sample_per_parent=50, simulate_rounds=100):
    """
    parent1: 第一个亲本的祖源列表
    parent2: 第二个亲本的祖源列表
    sample_per_parent: 从每个亲本采样的位点数,默认各50个
    simulate_rounds: 多轮模拟次数,次数越多结果越稳定
    返回所有模拟结果的合并统计
    """
    all_result = []
    for _ in range(simulate_rounds):
        # 从双方各采50%
        p1_sample = random.sample(parent1, sample_per_parent)
        p2_sample = random.sample(parent2, sample_per_parent)
        # 合并为孩子的祖源列表
        baby_sample = p1_sample + p2_sample
        all_result.extend(baby_sample)
    # 统计各祖源的占比
    result_df = pd.Series(all_result).value_counts(normalize=True).mul(100).round(2)
    return result_df

# 运行模拟,默认跑100轮
baby_ancestry = simulate_baby_ancestry(E, M)
print(baby_ancestry)

运行后会输出多轮模拟后孩子各祖源占比的平均结果,你可以调整simulate_rounds参数提高模拟次数,结果会更准确。

内容的提问来源于stack exchange,提问作者E_Sarousi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 19:54:01