R语言正则应用:按「数字.数字」格式拆分字符串生成新行
R拆分带编号文本DataFrame的解决方案
核心逻辑是修改正则匹配规则,直接匹配编号+对应内容的完整片段,用正向零宽断言限定匹配边界即可实现需求。
完整可运行代码
加载依赖包
library(tidyverse)
示例数据
dat <- data.frame(date= c("Sep2020", "Oct2020", "Nov2020", "Dec2020"), txt= c("1.1 What is the Constitution? 1.2 The original charter, which replaced the Articles of Confederation 1.3 hat all States would be equal. ", "4.4 What is the Bill of Rights? "4.5 The 9th and 10th amendments are general ", "5.1 in criminal prosecution to a speedy and public 5.2 War, three amendments were ratified (1865 5.3 13. The most recent amendment, the 27th, was", "6.2 the case of the proposed equal rights amendment, the Congress exten 6.3 but the proposed Amendment was never ratifie 6.4 tification deadline. The 38th State, Michig"))
核心处理代码
dat2 <- dat %>% # 匹配完整的编号+内容片段 mutate(txt = str_extract_all(txt, "\\d+\\.\\d+.*?(?=\\s*\\d+\\.\\d+|$)")) %>% # 拆分列表为多行 unnest_longer(txt) %>% # 可选:去除文本首尾多余空白,不需要可删除本行 mutate(txt = str_squish(txt))
正则规则说明
\\d+\\.\\d+.*?(?=\\s*\\d+\\.\\d+|$)的匹配逻辑:
\\d+\\.\\d+:匹配1位及以上数字+小数点+1位及以上数字的编号格式.*?:非贪婪匹配任意字符,尽可能短的匹配对应内容(?=\\s*\\d+\\.\\d+|$):正向零宽断言,匹配到「后面是可选空白+下一个编号」或者「字符串结尾」的位置就停止,不会占用下一个编号的内容
运行后得到的dat2结构和你预期的输出完全一致。
内容的提问来源于stack exchange,提问作者beeburt
相关产品推荐
相关产品推荐

