R语言技术问询:如何按指定引号类分隔符拆分单列文本为多列
解决R数据框单列评论拆分为多列的问题
问题分析
你之前误用了tidyr::separate_rows()——这个函数是把单列的多元素拆分成多行,和你要拆分成多列的需求完全不符;同时你写的分隔符sep ="",'\"语法错误,而且你的核心需求根本不是按分隔符拆分单条文本,而是把原本独立的每条评论从单列(列表或字符串列表形式)转成单独的列。
解决方案
情况1:comments列已是R列表格式
如果你的df$comments是R原生列表类型(每个列表元素对应一条评论),直接用tidyr::unnest_wider()将列表展开为多列:
library(tidyr) library(dplyr) # 模拟你的目标数据框 df <- tibble( comments = list( 'She really messed up bringing politics into climate change.... Climate change is a global problem. It should not be mixed with local conflicts.', "> She really messed up bringing politics into climate change\n\nEverything is politics, climate change more than anything. A shift in <em>policy</em> is what's climate activists are striving for. It's politics.", 'There are black-and-white political conflicts where taking sides could be safely done (like the Russian invasion of Ukraine, where Fridays for Future aligned with Ukraine), and there are gray-and-gray political topics where taking sides could only lead to division and polarization within the group (like the Israel-Palestine conflict).', 'but climate change has nothing to do with war. She is just splitting the movement', 'Israel-Palestine is really much less gray than what a lot of people are making it out to be. The Palestine people deserves to have their human rights upheld and protected just as any other people on the planet. Israel is very clearly in the process of carrying out ethnic cleansing and genocide against them. They should stop.', 'The war industry is the largest polluter. Working for peace and diplomacy is one of the most effective ways of reducing CO2e emissions.', "The human rights angle is one thing, but the territorial side of the dispute needs to be addressed as well (especially when it comes to Jerusalem). Both sides are currently taking the irredentist approach, which means the fighting won't stop until one of them secures exclusive and uncontested control of the area." ) ) # 转换为多列 new_df <- df %>% unnest_wider(comments, names_sep = "_")
运行后会生成comments_1到comments_7共7列,每列对应一条原始评论。
情况2:comments列是字符串形式的列表
如果你的df$comments是类似"['评论1','评论2',...]"的字符串(爬虫常见返回格式),需要先将其转换为R原生列表,再转成多列:
library(stringr) library(dplyr) library(tidyr) # 模拟字符串格式的列表列 df <- tibble( comments = "['She really messed up bringing politics into climate change.... Climate change is a global problem. It should not be mixed with local conflicts.', '> She really messed up bringing politics into climate change\n\nEverything is politics, climate change more than anything. A shift in <em>policy</em> is what's climate activists are striving for. It's politics.', 'There are black-and-white political conflicts where taking sides could be safely done (like the Russian invasion of Ukraine, where Fridays for Future aligned with Ukraine), and there are gray-and-gray political topics where taking sides could only lead to division and polarization within the group (like the Israel-Palestine conflict).', 'but climate change has nothing to do with war. She is just splitting the movement', 'Israel-Palestine is really much less gray than what a lot of people are making it out to be. The Palestine people deserves to have their human rights upheld and protected just as any other people on the planet. Israel is very clearly in the process of carrying out ethnic cleansing and genocide against them. They should stop.', 'The war industry is the largest polluter. Working for peace and diplomacy is one of the most effective ways of reducing CO2e emissions.', 'The human rights angle is one thing, but the territorial side of the dispute needs to be addressed as well (especially when it comes to Jerusalem). Both sides are currently taking the irredentist approach, which means the fighting won\\'t stop until one of them secures exclusive and uncontested control of the area.']" ) # 字符串转原生列表 df <- df %>% mutate(comments = str_remove_all(comments, "^\\[|\\]$") %>% # 移除首尾的[] str_split("',\\s*'") %>% # 按', '拆分字符串 map(function(x) str_remove_all(x, "^'|'$"))) # 移除每个元素的首尾引号 # 转换为多列 new_df <- df %>% unnest_wider(comments, names_sep = "_")
关键说明
unnest_wider()是专门处理列表列转多列的函数,names_sep参数用来控制列名的后缀格式。- 不要用
separate_rows(),它的作用是将单列拆分为多行,和你的需求完全相反。 - 如果后续需要拆分单条评论内部的文本,再针对具体列使用
separate()或正则表达式处理,但当前核心需求是将独立评论转为多列。
内容的提问来源于stack exchange,提问作者Daria
相关产品推荐
相关产品推荐

