You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中移除数据列的特定符号(括号、引号)

问题:移除数据列中的特定符号并提取首个流派内容

问题背景

处理Spotify 2022年全球流媒体数据时,需要清理genres列,移除剩余的括号、引号等符号,仅保留首个流派内容。当前genres列内容格式类似["dance pop",尝试用gsub移除符号未生效。

数据加载与预处理代码

library("dplyr")
library("stringr")
library("tidyverse")
library("ggplot2")

# 加载数据
spotiify_origional <- read.csv("https://raw.githubusercontent.com/info201a-au2022/project-group-1-section-aa/main/data/charts.csv")

# 筛选所需列
spotify_modify <- spotiify_origional %>% 
  select(name, country, date, position, streams, artists, genres = artist_genres)

# 提取2022年全球范围数据
spotify_2022_global <- spotify_modify %>% 
  filter(date >= "2022-01-01") %>% 
  filter(country == "global") %>% 
  arrange(date) %>% 
  group_by(streams)

用户尝试的清理代码

top_15 <- spotify_2022_global[order(spotify_2022_global$streams, decreasing = TRUE), ]
top_15 <- top_15[1:15,]
top_15$streams <- as.numeric(top_15$streams)

# 尝试拆分并清理符号
top_15 <- top_15 %>% 
  separate(genres, c("genres"), sep = ',')
top_15$genres<-gsub("]","",as.character(top_15$genres))

解决方案

问题源于目标符号是转义字符或属于类JSON数组结构,直接匹配普通字符无效,以下是三种可行方法:

方法1:正则表达式一次性清理

用正则匹配开头的["和结尾的],直接替换为空,同时保留首个流派:

top_15$genres <- gsub('^\\[\"|\\]$', '', as.character(top_15$genres))

如果需要精准提取首个流派,可直接截取["到第一个"之间的内容:

top_15$genres <- str_extract(as.character(top_15$genres), '(?<=\\[\").*?(?=\")')

方法2:JSON格式解析(更稳妥)

原始genres列是类JSON数组结构,用jsonlite包解析后取第一个元素:

library(jsonlite)
top_15$genres <- sapply(fromJSON(paste0("[", top_15$genres, "]")), function(x) x[1])

方法3:tidyverse工具链一步到位

结合str_remove_all和str_split完成清理与提取:

top_15 <- top_15 %>%
  mutate(genres = str_remove_all(as.character(genres), '\\[|\\]|\"') %>% 
           str_split(",") %>% 
           map_chr(~ .x[1]))

内容的提问来源于stack exchange,提问作者xvijerhe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 08:35:18