R中正则表达式无法匹配双换行,如何完整提取跨行问题?
问题解决:提取跨行的完整问题文本
问题背景
有一段以双换行\n\n分隔的文本,需要提取所有位于\n\n+数字.之后、下一个\n\n之前的完整问题(部分问题包含单换行跨行显示)。原使用R语言str_match_all搭配正则(?<=[\n\n]\d\.\s)(.*)(?=[\n\n])提取时,只能截取到问题中的单换行位置,无法获取完整跨行内容,需优化正则表达式。
原测试代码
txt <- "\n\n1. Is Albania classified as a least developed country by the United Nations?\n\n Currently, Albania is classified as a developing country by the United Nations Development\nProgramme, except for some UNDP programmes for which Albania is classified as a least developed\ncountry. Albania has applied for \"least-developed country\" status with UNDP and is awaiting a decision\non this matter.\n\nA.1(d) Division of Authority\n\n2. Please provide further information regarding the existing administrative structure in\nAlbania: ministries and associated agencies at the central level, local and regional authorities.\n\n Please see Attachment I for a complete description of Albania's administrative structure.\n\nA.2 Economic Situation and Policies\n\n3. An economic stabilization and reform programme was initiated in 1992. We would be\ninterested to learn more details about this programme.\n\n4. Is the \"comprehensive economic stabilization and reform programme (established in mid-\n1992) still in existence? If so, would you please describe what is still in effect.\n\n Albania's economic stabilization and reform programme is still in effect. With the achievement\nof basic stabilization, the emphasis of the programme has shifted to structural reforms. Nevertheless,\nthe government remains mindful of the importance of maintaining a stable macro-economic environment,\nand will continue its efforts in this direction.\n\n" questions <- str_match_all(txt,"(?<=[\n\n]\\d\\.\\s)(.*)(?=[\n\n])")
当前输出
[[1]] [,1] [1,] "Is Albania classified as a least developed country by the United Nations?" [2,] "Please provide further information regarding the existing administrative structure in" [3,] "An economic stabilization and reform programme was initiated in 1992. We would be" [4,] "Is the \"comprehensive economic stabilization and reform programme (established in mid-" [,2] [1,] "Is Albania classified as a least developed country by the United Nations?" [2,] "Please provide further information regarding the existing administrative structure in" [3,] "An economic stabilization and reform programme was initiated in 1992. We would be" [4,] "Is the \"comprehensive economic stabilization and reform programme (established in mid-"
预期输出
[[1]] [,1] [1,] "Is Albania classified as a least developed country by the United Nations?" [2,] "Please provide further information regarding the existing administrative structure in\nAlbania: ministries and associated agencies at the central level, local and regional authorities." [3,] "An economic stabilization and reform programme was initiated in 1992. We would be\ninterested to learn more details about this programme." [4,] "Is the \"comprehensive economic stabilization and reform programme (established in mid-\n1992) still in existence? If so, would you please describe what is still in effect." [,2] [1,] "Is Albania classified as a least developed country by the United Nations?" [2,] "Please provide further information regarding the existing administrative structure in\nAlbania: ministries and associated agencies at the central level, local and regional authorities." [3,] "An economic stabilization and reform programme was initiated in 1992. We would be\ninterested to learn more details about this programme." [4,] "Is the \"comprehensive economic stabilization and reform programme (established in mid-\n1992) still in existence? If so, would you please describe what is still in effect."
解决方案
原正则表达式中的.*默认不匹配换行符,因此只能捕获到问题中的第一行内容。优化后的正则需要匹配包括换行符在内的所有字符,直到遇到下一个双换行\n\n边界。
优化后的正则表达式:
(?<=[\n\n]\d\.\s)([\s\S]*?)(?=[\n\n])
[\s\S]:匹配任意字符(包括换行符,\s匹配空白字符,\S匹配非空白字符,两者结合覆盖所有字符)*?:非贪婪匹配,确保在遇到第一个\n\n时停止,避免匹配到后续内容
优化后的代码
txt <- "\n\n1. Is Albania classified as a least developed country by the United Nations?\n\n Currently, Albania is classified as a developing country by the United Nations Development\nProgramme, except for some UNDP programmes for which Albania is classified as a least developed\ncountry. Albania has applied for \"least-developed country\" status with UNDP and is awaiting a decision\non this matter.\n\nA.1(d) Division of Authority\n\n2. Please provide further information regarding the existing administrative structure in\nAlbania: ministries and associated agencies at the central level, local and regional authorities.\n\n Please see Attachment I for a complete description of Albania's administrative structure.\n\nA.2 Economic Situation and Policies\n\n3. An economic stabilization and reform programme was initiated in 1992. We would be\ninterested to learn more details about this programme.\n\n4. Is the \"comprehensive economic stabilization and reform programme (established in mid-\n1992) still in existence? If so, would you please describe what is still in effect.\n\n Albania's economic stabilization and reform programme is still in effect. With the achievement\nof basic stabilization, the emphasis of the programme has shifted to structural reforms. Nevertheless,\nthe government remains mindful of the importance of maintaining a stable macro-economic environment,\nand will continue its efforts in this direction.\n\n" questions <- str_match_all(txt,"(?<=[\n\n]\\d\\.\\s)([\\s\\S]*?)(?=[\n\n])")
运行后即可得到预期的完整跨行问题内容。
内容的提问来源于stack exchange,提问作者oly.f
相关产品推荐
相关产品推荐

