Python3从YAML替换JSON指定值遇TypeError及优化需求
Fixing the TypeError & Optimizing Bulk JSON Updates from YAML Config
Let's break down your problem and fix it step by step, plus optimize the solution for handling thousands of JSON files.
Why You're Getting the TypeError
The error TypeError: 'dict_keys' object does not support indexing comes from a Python 3 change: in Python 3, dict.keys() returns a view object (not a list like in Python 2), so you can't use [0] to access the first key directly. Your line:
yaml_config = { list(config.keys()[0]): list(config[config.keys()[0]]) for config in yaml_config }
is trying to index dict_keys, which isn't allowed in Python 3.
Fixed & Optimized Solution
First, we'll fix the YAML parsing logic, then switch to using glob to automatically find all JSON files (perfect for thousands of files instead of listing them manually).
Full Working Code
import json import yaml from glob import glob # Load and parse the YAML config correctly with open("master_conf.yaml", 'r') as f: yaml_config = yaml.safe_load(f) # Convert the list of single-key dicts into a flat dict (e.g., {"sourcesystem1": {...}, "sourcesystem2": {...}}) yaml_config = {next(iter(config.keys())): config[next(iter(config.keys()))] for config in yaml_config} # Get all JSON files in the current directory (adjust the pattern if needed, e.g., "./json_files/*.json") json_files = glob("*.json") for json_file in json_files: try: # Load the JSON file with open(json_file, "r") as f: sourcesystem_conf = json.load(f) # Get the source system name from the JSON sourcesystem = sourcesystem_conf.get("source") if not sourcesystem: print(f"Skipping {json_file}: no 'source' field found") continue # Update the JSON with values from YAML if the source exists if sourcesystem in yaml_config: # Option 1: Update all matching keys (including adding new ones from YAML) sourcesystem_conf.update(yaml_config[sourcesystem]) # Option 2: Only update fields that already exist in the JSON (uncomment below to use) # update_keys = set(yaml_config[sourcesystem].keys()) & set(sourcesystem_conf.keys()) # for key in update_keys: # sourcesystem_conf[key] = yaml_config[sourcesystem][key] # Write the updated JSON back with indentation with open(json_file, "w") as f: json.dump(sourcesystem_conf, f, indent=2) print(f"Updated {json_file} successfully") else: print(f"Skipping {json_file}: {sourcesystem} not found in YAML config") except Exception as e: print(f"Error processing {json_file}: {str(e)}")
Key Improvements
- Fixed YAML Parsing: Used
next(iter(config.keys()))to safely get the first (and only) key from each YAML entry, which works seamlessly in Python 3. - Bulk JSON Discovery:
glob("*.json")automatically finds all JSON files in the current directory. If your files are in a subfolder, useglob("./path/to/json/*.json")instead. - Error Handling: Added try/except blocks to catch issues with individual files without stopping the entire batch process.
- Flexible Updates: Included two update options—choose between adding new fields from YAML or only updating existing JSON fields, depending on your needs.
- User Feedback: Prints status messages so you can track which files are processed, skipped, or failed.
Additional Tips for Production Use
- Backup First: Before running on thousands of files, make a backup of your JSON directory, or modify the code to write updated files to a new folder instead of overwriting originals.
- Validate Configs: Add schema validation (e.g., with the
jsonschemalibrary) if you need to ensure YAML/JSON files follow a specific structure. - Performance: For extremely large numbers of files, you could use multiprocessing to speed up processing, but the sequential approach is sufficient and easier to debug for most cases.
内容的提问来源于stack exchange,提问作者user3764303
相关产品推荐
相关产品推荐

