如何在R语言DataFrame的指定列中去除小数部分?
Solution to Remove Decimal Part from
gene_id Column Got it, let's fix that gene_id column! You need to strip off everything after the decimal point while keeping the rest of your data frame unchanged. Here are a couple of easy, reliable methods using R:
Method 1: Using Tidyverse (dplyr + stringr)
Since you're already using read_csv() from the tidyverse, this method will fit right into your workflow:
library(tidyverse) # Load your data df <- read_csv("~/data.csv") # Clean the gene_id column df_cleaned <- df %>% mutate(gene_id = str_remove(gene_id, "\\..*"))
How this works:
str_remove(gene_id, "\\..*")targets the decimal point (\\.—we escape it because it's a special character in regex) and every character that comes after it (.*), then removes that entire segment.
Method 2: Base R (No Extra Packages)
If you prefer sticking to base R, you can use the sub() function for a concise solution:
# Load your data (base R alternative) df <- read.csv("~/data.csv") # Clean the gene_id column df$gene_id <- sub("\\..*", "", df$gene_id)
Or, if you want to use strsplit() instead for a more explicit split:
df$gene_id <- sapply(strsplit(as.character(df$gene_id), "\\."), "[", 1)
How these work:
sub("\\..*", "", df$gene_id)replaces the first occurrence of the "decimal point + all following characters" pattern with an empty string.strsplit()splits eachgene_idstring at the decimal point, thensapply()grabs the first element of each split result (the part before the decimal).
Verify the Result
Run head(df_cleaned) (or head(df) if you modified the data frame in-place) to confirm your gene_id column now looks like:
gene_id a b
abc100 30 70
abc101 40 80
abc102 50 90
abc103 60 100
内容的提问来源于stack exchange,提问作者user7755336
相关产品推荐
相关产品推荐

