如何在RethinkDB中展开合并层级文档结构的嵌套数组得到单个数组
问题描述
我有一个嵌套文档结构,已经可以通过pluck筛选出相关部分:
是否有优雅的方案可以将最底层的所有条目合并为单个数组?
预期结果(条目特意保留重复值):
[ '3425b91f-f019-4db3-ad56-c336bf55279b', '3d07946e-183d-4992-9acd-676f5122e1b1', '3425b91f-f019-4db3-ad56-c336bf55279b', '3d07946e-183d-4992-9acd-676f5122e1b1', '2cd652a6-4dcd-4920-9592-d4cdc5a034bf', '70fe1812-e1de-447b-ac4f-d89fead4756d', '2cd652a6-4dcd-4920-9592-d4cdc5a034bf', '70fe1812-e1de-447b-ac4f-d89fead4756d' ]
我尝试使用如下写法:
r.table('periods')['regions']['sites']['plants']['product']['process']['technologies'].run()
但返回报错 "Cannot perform bracket on a sequence of sequences"。
=> 是否存在替代运算符,可以在每一步获取合并后的序列,而非「序列的序列」?
期望实现类似如下写法的功能:
r.table('periods').unwind('regions.sites.plants.product.process.technologies')
以下是生成示例数据的Python代码:
from rethinkdb import RethinkDB r = RethinkDB() r.connect({}).repl() r.table_create("periods") def uniqueid(): return r.uuid().run() periodid_first = uniqueid() periodid_second = uniqueid() companyid_2000 = uniqueid() companyid_2001 = uniqueid() technologyid_2000_first = uniqueid() technologyid_2000_second = uniqueid() technologyid_2001_first = uniqueid() technologyid_2001_second = uniqueid() energy_carrierid_2000_first = uniqueid() energy_carrierid_2000_second = uniqueid() energy_carrierid_2001_first = uniqueid() energy_carrierid_2001_second = uniqueid() periods = [ { 'id': periodid_first, 'start': 2000, 'end': 2000, # 'sub_periods': [], 'regions': [ { 'id': 'DE', # 'sub_regions': [], 'sites': [ { 'id': 'first_site_in_germany', 'company': companyid_2000, # => verweist auf periods => companies 'plants': [ { 'id': 'qux', 'product': { 'id': 'Ammoniak', 'process': { 'id': 'SMR+HB', 'technologies': [ technologyid_2000_first, # => verweist auf periods => technologies technologyid_2000_second ] } } } ] } ] }, { 'id': 'FR', # 'sub_regions': [], 'sites': [ { 'id': 'first_site_in_france', 'company': companyid_2000, # => verweist auf periods => companies 'plants': [ { 'id': 'qux', 'product': { 'id': 'Ammoniak', 'process': { 'id': 'SMR+HB', 'technologies': [ technologyid_2000_first, # => verweist auf periods => technologies technologyid_2000_second ] } } } ] } ] } ], 'companies': [ { 'id': companyid_2000, 'name': 'international_company' } ], 'technologies': [ { 'id': technologyid_2000_first, 'name': 'SMR', 'specific_cost_per_year': 123, 'specific_energy_consumptions': [ { 'energy_carrier': energy_carrierid_2000_first, 'specific_consumption': 5555 }, # => verweist auf periods => energy_carriers { 'energy_carrier': energy_carrierid_2000_second, 'energy_consumption': 2333 } ] }, { 'id': technologyid_2000_second, 'name': 'HB', 'specific_cost_per_year': 1234, 'specific_energy_consumptions': [ { 'energy_carrier': energy_carrierid_2000_first, 'specific_consumption': 555 }, # => verweist auf periods => energy_carriers { 'energy_carrier': energy_carrierid_2000_second, 'energy_consumption': 233 } ] } ], 'energy_carriers': [ { 'id': energy_carrierid_2000_first, 'name': 'oil', 'group': 'fuel' }, { 'id': energy_carrierid_2000_second, 'name': 'gas', 'group': 'fuel' }, { 'id': uniqueid(), 'name': 'conventional', 'group': 'electricity' }, { 'id': uniqueid(), 'name': 'green', 'group': 'electricity' } ], 'networks': [ { 'id': uniqueid(), 'name': 'gas', 'sub_networks': [], 'pipelines': [ ] }, { 'id': uniqueid(), 'name': 'gas', 'sub_networks': [], 'pipelines': [ ] } ] }, { 'id': periodid_second, 'start': 2001, 'end': 2001, # 'sub_periods': [], 'regions': [ { 'id': 'DE', # 'sub_regions': [], 'sites': [ { 'id': 'first_site_in_germany', 'company': companyid_2001, # => verweist auf periods => companies 'plants': [ { 'id': 'qux', 'product': { 'id': 'Ammoniak', 'process': { 'id': 'SMR+HB', 'technologies': [ technologyid_2001_first, # => verweist auf periods => technologies technologyid_2001_second ] } } } ] } ] }, { 'id': 'FR', # 'sub_regions': [], 'sites': [ { 'id': 'first_site_in_france', 'company': companyid_2001, # => verweist auf periods => companies 'plants': [ { 'id': 'qux', 'product': { 'id': 'Ammoniak', 'process': { 'id': 'SMR+HB', 'technologies': [ technologyid_2001_first, # => verweist auf periods => technologies technologyid_2001_second ] } } } ] } ] } ], 'companies': [ { 'id': companyid_2001, 'name': 'international_company' } ], 'technologies': [ { 'id': technologyid_2001_first, 'name': 'SMR', 'specific_cost_per_year': 123, 'specific_energy_consumptions': [ { 'energy_carrier': energy_carrierid_2001_first, 'specific_consumption': 5555 }, # => verweist auf periods => energy_carriers { 'energy_carrier': energy_carrierid_2001_second, 'energy_consumption': 2333 } ] }, { 'id': technologyid_2001_second, 'name': 'HB', 'specific_cost_per_year': 1234, 'specific_energy_consumptions': [ { 'energy_carrier': energy_carrierid_2001_first, 'specific_consumption': 555 }, # => verweist auf periods => energy_carriers { 'energy_carrier': energy_carrierid_2001_second, 'energy_consumption': 233 } ] } ], 'energy_carrieriers': [ { 'id': energy_carrierid_2001_first, 'name': 'oil', 'group': 'fuel' }, { 'id': energy_carrierid_2001_second, 'name': 'gas', 'group': 'fuel' }, { 'id': uniqueid(), 'name': 'conventional', 'group': 'electricity' }, { 'id': uniqueid(), 'name': 'green', 'group': 'electricity' } ], 'networks': [ { 'id': uniqueid(), 'name': 'gas', 'sub_networks': [], 'pipelines': [ ] }, { 'id': uniqueid(), 'name': 'gas', 'sub_networks': [], 'pipelines': [ ] } ] } ] r.table('periods') \ .insert(periods) \ .run()
解决方案
RethinkDB中可以使用concat_map运算符展开嵌套数组,每一层数组都通过concat_map拍平为一维序列后再继续处理下一层即可。
基础写法
r.table('periods')\ .concat_map(lambda period: period['regions'])\ .concat_map(lambda region: region['sites'])\ .concat_map(lambda site: site['plants'])\ .concat_map(lambda plant: plant['product']['process']['technologies'])\ .run()
该写法会保留所有重复的条目,完全符合预期输出要求。
封装为通用工具函数
如果需要频繁使用类似的嵌套路径展开功能,可以封装一个通用函数实现你期望的unwind调用效果:
def unwind(rql_obj, nested_path): fields = nested_path.split('.') current = rql_obj for field in fields: current = current.concat_map(lambda item: item[field]) return current # 调用方式和你期望的写法完全一致 unwind(r.table('periods'), 'regions.sites.plants.product.process.technologies').run()
内容的提问来源于stack exchange,提问作者Stefan
相关产品推荐
相关产品推荐

