You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求生成APOL1基因型汇总表并导出至同目录CSV文件

Got it, let's walk through how to create your desired Summary_table and save it to a CSV file using pandas. Here's a step-by-step solution tailored to your requirements:

Step 1: Define Mapping Rules

First, we'll create dictionaries to map your raw genotypes to both standard genotype names and risk allele count categories, as you specified:

import pandas as pd

# Map raw genotypes to standard display names
genotype_to_standard = {
    "G0/G0": "G0/G0",
    "G0/G2": "G0/G2",
    "G1^{GM}/G0": "G1/G0",
    "G1^{G+}/G0": "G1/G0",
    "G1^{GM}/G1^{GM}": "G1/G1",
    "G1^{GM}/G2": "G1/G2",
    "G2/G2": "G2/G2"
}

# Map raw genotypes to risk allele count categories
genotype_to_risk_count = {
    "G0/G0": 0,
    "G0/G2": 1,
    "G1^{GM}/G0": 1,
    "G1^{G+}/G0": 1,
    "G1^{GM}/G1^{GM}": 2,
    "G1^{GM}/G2": 2,
    "G2/G2": 2
}

Step 2: Transform the Original Dataframe

Apply these mappings to your Final_APOL1 dataframe to add the required columns:

# Assuming your genotype column is named "Final genotype of APOL1"
Final_APOL1["Standard Genotype"] = Final_APOL1["Final genotype of APOL1"].map(genotype_to_standard)
Final_APOL1["Risk Allele Count"] = Final_APOL1["Final genotype of APOL1"].map(genotype_to_risk_count)

# Optional: Drop rows with unrecognized genotypes (if any)
Final_APOL1 = Final_APOL1.dropna(subset=["Standard Genotype", "Risk Allele Count"])

Step 3: Generate the Summary Table

Calculate counts and frequencies grouped by risk allele category and standard genotype:

total_samples = len(Final_APOL1)

# Group and calculate counts
summary_df = Final_APOL1.groupby(["Risk Allele Count", "Standard Genotype"]).size().reset_index(name="Sample Count")

# Calculate individual genotype frequency (percentage of total population)
summary_df["Frequency (%)"] = round((summary_df["Sample Count"] / total_samples) * 100, 2)

# Add total frequency per risk allele category
risk_total = summary_df.groupby("Risk Allele Count")["Frequency (%)"].sum().reset_index(name="Total Category Frequency (%)")
summary_df = summary_df.merge(risk_total, on="Risk Allele Count")

# Replace numeric count with descriptive category names
summary_df["Risk Allele Category"] = summary_df["Risk Allele Count"].map({
    0: "no APOL1 Risk Alleles",
    1: "1 APOL1 Risk Alleles",
    2: "2 APOL1 Risk Alleles"
})

# Reorder columns for readability
summary_df = summary_df[["Risk Allele Category", "Standard Genotype", "Sample Count", "Frequency (%)", "Total Category Frequency (%)"]]

Step 4: Save to CSV

Yes, you can easily save this summary table to your current directory as a standalone CSV file:

summary_df.to_csv("Summary_table.csv", index=False, encoding="utf-8")

Data Structure Explanation

The output Summary_table.csv will have these columns:

  • Risk Allele Category: Descriptive group name (0/1/2 risk alleles)
  • Standard Genotype: Simplified genotype name (e.g., G1/G2 for G1^{GM}/G2)
  • Sample Count: Number of individuals with this genotype
  • Frequency (%): Percentage of the total population with this genotype
  • Total Category Frequency (%): Combined percentage of all genotypes in the same risk allele group

内容的提问来源于stack exchange,提问作者user12633410

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 11:57:44