请求生成APOL1基因型汇总表并导出至同目录CSV文件
Got it, let's walk through how to create your desired Summary_table and save it to a CSV file using pandas. Here's a step-by-step solution tailored to your requirements:
Step 1: Define Mapping Rules
First, we'll create dictionaries to map your raw genotypes to both standard genotype names and risk allele count categories, as you specified:
import pandas as pd # Map raw genotypes to standard display names genotype_to_standard = { "G0/G0": "G0/G0", "G0/G2": "G0/G2", "G1^{GM}/G0": "G1/G0", "G1^{G+}/G0": "G1/G0", "G1^{GM}/G1^{GM}": "G1/G1", "G1^{GM}/G2": "G1/G2", "G2/G2": "G2/G2" } # Map raw genotypes to risk allele count categories genotype_to_risk_count = { "G0/G0": 0, "G0/G2": 1, "G1^{GM}/G0": 1, "G1^{G+}/G0": 1, "G1^{GM}/G1^{GM}": 2, "G1^{GM}/G2": 2, "G2/G2": 2 }
Step 2: Transform the Original Dataframe
Apply these mappings to your Final_APOL1 dataframe to add the required columns:
# Assuming your genotype column is named "Final genotype of APOL1" Final_APOL1["Standard Genotype"] = Final_APOL1["Final genotype of APOL1"].map(genotype_to_standard) Final_APOL1["Risk Allele Count"] = Final_APOL1["Final genotype of APOL1"].map(genotype_to_risk_count) # Optional: Drop rows with unrecognized genotypes (if any) Final_APOL1 = Final_APOL1.dropna(subset=["Standard Genotype", "Risk Allele Count"])
Step 3: Generate the Summary Table
Calculate counts and frequencies grouped by risk allele category and standard genotype:
total_samples = len(Final_APOL1) # Group and calculate counts summary_df = Final_APOL1.groupby(["Risk Allele Count", "Standard Genotype"]).size().reset_index(name="Sample Count") # Calculate individual genotype frequency (percentage of total population) summary_df["Frequency (%)"] = round((summary_df["Sample Count"] / total_samples) * 100, 2) # Add total frequency per risk allele category risk_total = summary_df.groupby("Risk Allele Count")["Frequency (%)"].sum().reset_index(name="Total Category Frequency (%)") summary_df = summary_df.merge(risk_total, on="Risk Allele Count") # Replace numeric count with descriptive category names summary_df["Risk Allele Category"] = summary_df["Risk Allele Count"].map({ 0: "no APOL1 Risk Alleles", 1: "1 APOL1 Risk Alleles", 2: "2 APOL1 Risk Alleles" }) # Reorder columns for readability summary_df = summary_df[["Risk Allele Category", "Standard Genotype", "Sample Count", "Frequency (%)", "Total Category Frequency (%)"]]
Step 4: Save to CSV
Yes, you can easily save this summary table to your current directory as a standalone CSV file:
summary_df.to_csv("Summary_table.csv", index=False, encoding="utf-8")
Data Structure Explanation
The output Summary_table.csv will have these columns:
- Risk Allele Category: Descriptive group name (0/1/2 risk alleles)
- Standard Genotype: Simplified genotype name (e.g.,
G1/G2forG1^{GM}/G2) - Sample Count: Number of individuals with this genotype
- Frequency (%): Percentage of the total population with this genotype
- Total Category Frequency (%): Combined percentage of all genotypes in the same risk allele group
内容的提问来源于stack exchange,提问作者user12633410
相关产品推荐
相关产品推荐

