如何将Python列表转为字典并去重整合分类标签?
问题描述
将Python列表转换为字典以便输出为JSON时遇到两个问题:
- 循环处理列表时,每个标签生成独立的包含
data_tags的字典,未整合到同一列表 - 重复标签(如Torpedo)被重复添加
需要实现:将所有不重复的分类标签整合到同一个data_tags列表中。
现有代码片段
for item in class_resources: tag = item print(tag)
owner_name = output.get('provider', {}).get('name') or 'NBN Atlas' owner_url = f"https://registry.nbnatlas.org/public/show/{uid}" if uid is not None else "https://registry.nbnatlas.org/datasets" owner_api_url = output.get('provider', {}).get( 'uri') or 'https://registry.nbnatlas.org/ws/dataResource' output_dict = { 'data_providers': [ { 'name': 'NBN Atlas', 'url': 'https://registry.nbnatlas.org/datasets', 'api_url': 'https://registry.nbnatlas.org/ws/dataResource', 'type': 'host' }, { 'name': owner_name, 'url': owner_url, 'api_url': owner_api_url, 'type': 'owner' } ], 'data_tags': [{ 'type': 'classification', 'tag': tag, 'tag_lower': tag.lower() }] }
当前输出
{ 'data_tags': [ { 'tag': 'Pisces', 'tag_lower': 'pisces', 'type': 'classification'}]} { 'data_tags': [ { 'tag': 'Animalia', 'tag_lower': 'animalia', 'type': 'classification'}]} ... { 'data_tags': [ { 'tag': 'Torpedo', 'tag_lower': 'torpedo', 'type': 'classification'}]} { 'data_tags': [ { 'tag': 'Torpedo', 'tag_lower': 'torpedo', 'type': 'classification'}]}
期望输出
"data_tags": [ { "type": "classification", "tag": "Pisces", "tag_lower": "pisces" }, { "type": "classification", "tag": "Animalia", "tag_lower": "animalia" }, ... ],
解决方案
步骤说明
- 去重处理:先对
class_resources中的标签去重,避免重复添加。如果需要保留原顺序,可使用dict.fromkeys()(Python 3.7+字典有序);不需要顺序的话直接转集合即可。 - 构建统一的data_tags列表:遍历去重后的标签,生成每个标签对应的字典,收集到同一个列表中。
- 整合到output_dict:将生成的
data_tags列表放入最终的输出字典,而非每次循环生成新字典。
修改后的代码
# 对标签去重,保留原顺序(Python 3.7+) unique_tags = list(dict.fromkeys(class_resources)) # 构建data_tags列表 data_tags = [] for tag in unique_tags: data_tags.append({ 'type': 'classification', 'tag': tag, 'tag_lower': tag.lower() }) # 生成最终输出字典 owner_name = output.get('provider', {}).get('name') or 'NBN Atlas' owner_url = f"https://registry.nbnatlas.org/public/show/{uid}" if uid is not None else "https://registry.nbnatlas.org/datasets" owner_api_url = output.get('provider', {}).get( 'uri') or 'https://registry.nbnatlas.org/ws/dataResource' output_dict = { 'data_providers': [ { 'name': 'NBN Atlas', 'url': 'https://registry.nbnatlas.org/datasets', 'api_url': 'https://registry.nbnatlas.org/ws/dataResource', 'type': 'host' }, { 'name': owner_name, 'url': owner_url, 'api_url': owner_api_url, 'type': 'owner' } ], 'data_tags': data_tags } # 输出JSON格式结果 import json print(json.dumps(output_dict, indent=2))
说明
- 使用
dict.fromkeys(class_resources)去重能保留标签在原列表中的出现顺序,若无需顺序,可简化为unique_tags = list(set(class_resources))。 - 把
data_tags的构建放在循环外,确保所有标签都被整合到同一个列表中,避免每次循环生成独立的输出字典。
内容的提问来源于stack exchange,提问作者h1m aga1n
相关产品推荐
相关产品推荐

