PCRE2正则能否在单次替换中处理捕获组内的多引用内容?
问题:CSV转JSON时如何单次正则替换捕获所有引用内容?
我需要将Excel导出的CSV调研结果转换为网站地图用的JSON,每条记录包含多条以Story xx: 开头、换行结尾的引用。
示例CSV记录:
1,Alethia Tanner Park,227 Harry Thomas Way NE,Washington,DC,20002,3,"Story 1: First quote of someone commenting on this location. Story 2: Second quote from a resident that likes this place. Story 3: Third quote about this park in the city. ",38.91143761,-77.00165721
期望JSON格式:
{ "latitudeLongitude": [38.91143761, -77.00165721], "displayName": "Alethia Tanner Park", "quote": "“First quote of someone commenting on this location.”<hr />“Second quote from a resident that likes this place.”<hr />“Third quote about this park in the city.”<hr />", "numberOfCitations": 3 },
当前问题:
我使用PCRE2引擎的regex101工具,当前正则只能捕获最后一条引用,替换后仅输出最后一条的处理结果;尝试用(“${quote}”<hr />)+也无效。想确认是否能通过单次替换完成所有引用的捕获和处理——分步处理可以实现,但好奇单次替换的可行性。
解决方案
核心原因
你使用的重复捕获组(?'quote'...)在PCRE2中只会保留最后一次匹配的内容,所以${quote}只能拿到最后一条引用,这是正则引擎的固有特性:重复捕获的分组不会存储所有匹配结果。
单次替换实现方法(PCRE2)
要在单次替换中完成需求,需要结合回调函数替换(PCRE2原生支持,regex101也适配此功能),具体步骤如下:
- 匹配整条记录的正则:
^(?'groupId'\d{1,3}),(?'displayName'[^,]+),(?'street'\?\?\?|\d+ [\w ]+),(?'city'\?\?\?|[A-Za-z ]+),(?'state'\?\?\?|[A-Z]{2}),(?'zipCode'\?\?\?|\d{5}),(?'numberOfCitations'\d{1,2}),"(?'quotes'(?:Story \d{1,2}: .+?(?=\nStory|\n"))+)",(?'latitude'38\.\d{4,8}),(?'longitude'-7[67]\.\d{4,8})$
该正则将所有引用内容整体捕获到quotes分组,不再拆分单个quote。
- 回调函数处理逻辑:
在替换时,通过回调对捕获到的引用内容做二次转换,再组装成目标JSON:
preg_replace_callback( '/^(?'groupId'\d{1,3}),(?'displayName'[^,]+),(?'street'\?\?\?|\d+ [\w ]+),(?'city'\?\?\?|[A-Za-z ]+),(?'state'\?\?\?|[A-Z]{2}),(?'zipCode'\?\?\?|\d{5}),(?'numberOfCitations'\d{1,2}),"(?'quotes'(?:Story \d{1,2}: .+?(?=\nStory|\n"))+)",(?'latitude'38\.\d{4,8}),(?'longitude'-7[67]\.\d{4,8})$/', function($matches) { // 把每条Story格式转换为带HTML引号和水平线的内容 $processedQuotes = preg_replace( '/Story \d{1,2}: (.+?)(?=\nStory|\n")/', '“$1”<hr />', $matches['quotes'] ); // 组装最终JSON字符串 return "{ \"latitudeLongitude\": [{$matches['latitude']}, {$matches['longitude']}], \"displayName\": \"{$matches['displayName']}\", \"quote\": \"{$processedQuotes}\", \"numberOfCitations\": {$matches['numberOfCitations']} },"; }, $csvContent );
纯正则模拟单次处理(regex101环境)
如果不想用回调函数,可通过连续两次全局替换实现等效的单次批量处理:
- 第一步:全局替换所有引用行的格式
正则:/Story \d{1,2}: (.+?)(?=\n)/
替换:“$1”<hr /> - 第二步:全局替换每条CSV记录为JSON结构
正则:/^(?'groupId'\d{1,3}),(?'displayName'[^,]+),(?'street'\?\?\?|\d+ [\w ]+),(?'city'\?\?\?|[A-Za-z ]+),(?'state'\?\?\?|[A-Z]{2}),(?'zipCode'\?\?\?|\d{5}),(?'numberOfCitations'\d{1,2}),"(?'quotes'[^"]+)",(?'latitude'38\.\d{4,8}),(?'longitude'-7[67]\.\d{4,8})$/' 替换:{\n\t"latitudeLongitude": [${latitude}, ${longitude}], \n\t"displayName": "${displayName}", \n\t"quote": "${quotes}", \n\t"numberOfCitations": ${numberOfCitations}\n},`
内容的提问来源于stack exchange,提问作者PeterV
相关产品推荐
相关产品推荐

