如何使用预查正则表达式拆分指定JSON格式字符串?
拆分JSON对象字符串的预查正则方案
Got it, let's walk through how to split your concatenated JSON cost records using lookahead regex. Your input is a string of back-to-back AWS Cost Explorer-style objects (missing the outer array brackets), so we need a way to split exactly between each complete object without breaking the structure.
核心正则表达式
Here's the regex that will do the trick:
}(?=\s*{"TimePeriod":|$)
正则逻辑拆解
}: 匹配单个成本记录对象的闭合大括号(?=\s*{"TimePeriod":|$): 这是一个正向预查,确保这个闭合大括号后面要么是:- 可选的空白字符(
\s*)加上新记录的起始标记({"TimePeriod":),要么是 - 字符串结尾(
$),用来处理最后一个对象
- 可选的空白字符(
示例代码(Python)
Let's test this with your input string (I'll fill in the truncated last object for demonstration):
import re input_str = '''{"TimePeriod": {"Start": "2017-03-01", "End": "2017-04-01"}, "Total": {"UnblendedCost": {"Amount": "2942.25119998", "Unit": "USD"}, "UsageQuantity": {"Amount": "20835", "Unit": "Hrs"}}, "Groups": [], "Estimated": false}, {"TimePeriod": {"Start": "2017-04-01", "End": "2017-05-01"}, "Total": {"UnblendedCost": {"Amount": "2982.62609983", "Unit": "USD"}, "UsageQuantity": {"Amount": "21049", "Unit": "Hrs"}}, "Groups": [], "Estimated": false}, {"TimePeriod": {"Start": "2017-05-01", "End": "2017-06-01"}, "Total": {"UnblendedCost": {"Amount": "3010.45", "Unit": "USD"}, "UsageQuantity": {"Amount": "21200", "Unit": "Hrs"}}, "Groups": [], "Estimated": false}''' # 使用预查正则拆分字符串 split_records = re.split(r'}(?=\s*{"TimePeriod":|$)', input_str) # 修复拆分后丢失的闭合大括号 split_records = [record + '}' if not record.endswith('}') else record for record in split_records] # 输出拆分结果 for idx, record in enumerate(split_records, 1): print(f"--- 记录 {idx} ---") print(record)
注意事项
- 如果你的输入包含截断的不完整记录(比如你原字符串里的
{"TimePeriod": {"Start": "2017-05-01",...),正则会把它当作最后一个元素保留。你可以额外添加逻辑过滤掉这类不完整对象。 - 这个正则假设所有记录都以
{"TimePeriod":开头,如果你的记录结构有变化,需要调整预查里的匹配模式。
内容的提问来源于stack exchange,提问作者Forrest
相关产品推荐
相关产品推荐

