如何修改Java正则表达式以将多行笔记内容捕获到单个捕获组?
问题描述
在我的应用中,提取的笔记格式如下:
**************************** Example noteText krishna 2023-05-16T11:29:59Z **************************** ***************************** Example notes for testing Extension of the same notes 2023-05-16T11:36:03Z ******************************
需要将这些笔记转换为指定格式的JSON,预期结果如下:
[ { "timestamp": "2023-05-16T11:29:59Z", "userid": "krishna", "content": "Example noteText" }, { "timestamp": "2023-05-16T11:36:03Z", "userid": "krishna", "content": "Example notes for testing Extension of the same notes" } ]
已尝试的解决方案
我当前用Java正则表达式实现,但无法把多行笔记内容捕获到单个分组里,导致一条多行笔记被拆成多条记录。
代码
Pattern notesPattern = Pattern.compile( "(?:\\*{25}\\r\\n)" + "(?<date>[\\w\\s:]+?)" + "\\r\\n" + "(?<userid>\\w+?)" + "\\r\\n" + "(?:\\*{25}\\r\\n)" + "(?<content>.+)" + // Note content is captured here "(?:\\r\\n\\r\\n)?");
当前输出
[ { "timestamp": "2023-05-16T11:29:59Z", "userid": "krishna", "content": "Example noteText" }, { "timestamp": "2023-05-16T11:36:03Z", "userid": "krishna", "content": "Example notes for testing " }, { "timestamp": null, "userid": null, "content": "\\n\\nExtension of the same notes" } ]
解决方案
问题根源是原正则的(?<content>.+)默认不匹配换行符,且未限定内容的终止边界。调整后的正则会启用多行匹配模式,并通过负前瞻锁定内容的结束位置,确保完整捕获多行笔记内容:
Pattern notesPattern = Pattern.compile( "(?:\\*{25,30}\\r?\\n)" + // 适配不同长度的星号分隔线 "(?<content>(?:(?!\\r?\\n\\w+\\s+\\d{4}-\\d{2}-\\d{2}T).)+)" + // 捕获到用户+时间戳行前的所有内容 "\\r?\\n" + "(?<userid>\\w+)" + "\\s+" + "(?<timestamp>\\d{4}-\\d{2}-\\d{2}T\\d{2}:\\d{2}:\\d{2}Z)" + "\\r?\\n\\s*" + "(?:\\*{25,30})", Pattern.DOTALL // 让.匹配换行符,支持多行内容捕获 );
核心调整点:
- 启用
Pattern.DOTALL:让正则中的.可以匹配换行符,实现多行内容的捕获。 - 负前瞻限定内容边界:
(?!\\r?\\n\\w+\\s+\\d{4}-\\d{2}-\\d{2}T)确保内容部分只匹配到用户ID和时间戳行之前的所有文本,避免拆分记录。 - 兼容分隔线长度:用
\\*{25,30}适配示例中不同数量的星号分隔符。 - 兼容换行格式:
\\r?\\n同时支持Windows和Unix风格的换行符。
捕获到内容后,只需将内容中的换行符替换为空格,就能得到预期的JSON结构。
内容的提问来源于stack exchange,提问作者krishna lakshmi
相关产品推荐
相关产品推荐

