如何从Sample列提取数据创建tissue、plant、stage新列?
Absolutely—your sample IDs follow a clear, consistent pattern: [tissue code][plant number].[stage number], which makes it straightforward to extract those three components into new columns. Below are step-by-step solutions using two common data analysis tools: Python's Pandas and R.
Solution 1: Using Python Pandas
Assuming your data is stored in a Pandas DataFrame, here's how to split the sample column:
- Set up your data:
import pandas as pd # Example data (replace with your actual dataset) data = { 'sample': ['P10.1', 'P11.2', 'S1.1', 'S3.3'], # Include your existing measurement columns here (P10.1, P11.2, etc.) } df = pd.DataFrame(data)
- Extract components with regex:
Usestr.extract()to pull out each part using a regular expression tailored to your ID structure:
# Regex breakdown: # ^([A-Za-z]) → Captures the first letter (tissue code) # (\d+) → Captures digits before the dot (plant number) # \. → Matches the dot separator # (\d+)$ → Captures digits after the dot (stage number) df[['tissue', 'plant', 'stage']] = df['sample'].str.extract(r'^([A-Za-z])(\d+)\.(\d+)$') # Optional: Convert plant/stage to numeric types for easier analysis df['plant'] = df['plant'].astype(int) df['stage'] = df['stage'].astype(int)
- Final Result:
Your DataFrame will now include the three new columns, exactly matching your expected output.
Solution 2: Using R
If you prefer working in R, here are two clean methods to split the sample column:
Option A: Using tidyr::separate()
This method uses regex to split the column at logical points:
library(tidyr) # Example data df <- data.frame( sample = c("P10.1", "P11.2", "S1.1", "S3.3"), # Add your existing measurement columns here ) # Split the sample column into three parts df <- df %>% separate( sample, into = c("tissue", "plant", "stage"), sep = "(?<=[A-Za-z])(?=\\d)|\\." # Splits after the letter and at the dot ) # Convert plant/stage to numeric df$plant <- as.integer(df$plant) df$stage <- as.integer(df$stage)
Option B: Using stringr for direct regex extraction
If you want more control over each component:
library(stringr) df$tissue <- str_extract(df$sample, "^[A-Za-z]") df$plant <- as.integer(str_extract(df$sample, "(?<=^[A-Za-z])\\d+(?=\\.)")) df$stage <- as.integer(str_extract(df$sample, "(?<=\\.)\\d+$"))
Both approaches will populate your tissue, plant, and stage columns perfectly from the sample IDs. Just adjust the code to match your actual dataset's structure.
内容的提问来源于stack exchange,提问作者Lauren Maynard

