寻求更简洁的Python字典验证内置实现方案
Hey there! Your existing code works for checking key existence and value type matches, but the path-generation helper functions make it tough to read and maintain. Let's walk through a few more elegant solutions—starting with pure built-in Python, then a popular library for more robust use cases.
Option 1: Recursive Validation (Pure Built-In, Readable)
We can use a recursive function to traverse both your target dictionary and the validation schema directly. This keeps the logic linear and easy to follow, no need to generate separate path lists:
check_against = {'a': str, 'b': {'c': int, 'd': int}} a = {'a': 1, 'c': 1} def validate_dict(obj, schema, current_path=""): for key, expected_type in schema.items(): # Build a human-readable path string full_path = f"{current_path}.{key}" if current_path else key path_list_str = "['" + "', '".join(full_path.split('.')) + "']" # Check if the key exists in the target dictionary if key not in obj: print(f"Missing key at path \"{path_list_str}\"") continue # If the expected type is a nested dict, recurse to validate the substructure if isinstance(expected_type, dict): validate_dict(obj[key], expected_type, full_path) # Otherwise, check if the value matches the expected type else: actual_type = type(obj[key]) if not isinstance(obj[key], expected_type): print(f"Value at path \"{path_list_str}\" should be of type \"{expected_type}\" but got {actual_type}") # Run the validation validate_dict(a, check_against)
This will output exactly what your original code does, but with much clearer logic:
Value at path "['a']" should be of type "<class 'str'>" but got <class 'int'> Missing key at path "['b', 'c']" Missing key at path "['b', 'd']"
Option 2: Use Dataclasses (Python 3.7+ Built-In)
If your data structure is well-defined, Python's built-in dataclasses paired with type hints can make validation more structured. This turns your schema into explicit class definitions, which are self-documenting:
from dataclasses import dataclass # Define your schema as dataclasses @dataclass class SubSchema: c: int d: int @dataclass class MainSchema: a: str b: SubSchema def validate_with_dataclass(obj, schema_cls, current_path=""): for field_name, field in schema_cls.__dataclass_fields__.items(): full_path = f"{current_path}.{field_name}" if current_path else field_name path_list_str = "['" + "', '".join(full_path.split('.')) + "']" if field_name not in obj: print(f"Missing key at path \"{path_list_str}\"") continue # Recurse if the field type is another dataclass if hasattr(field.type, '__dataclass_fields__'): validate_with_dataclass(obj[field_name], field.type, full_path) else: actual_type = type(obj[field_name]) if not isinstance(obj[field_name], field.type): print(f"Value at path \"{path_list_str}\" should be of type \"{field.type}\" but got {actual_type}") # Run the validation validate_with_dataclass(a, MainSchema)
This approach is great for scenarios where your data follows a strict structure—you get the added benefit of type hints for your codebase too.
Option 3: Pydantic (Industry-Standard, Non-Built-In)
If you need more robust validation (like default values, format checks, or automatic error messaging), Pydantic is the go-to library. It handles most of the heavy lifting for you, with minimal code:
from pydantic import BaseModel, ValidationError # Define your schema with Pydantic models class SubSchema(BaseModel): c: int d: int class MainSchema(BaseModel): a: str b: SubSchema try: # Try to parse the dictionary into the Pydantic model MainSchema(**a) except ValidationError as e: # Format the errors to match your original output style for error in e.errors(): path_list_str = "['" + "', '".join(map(str, error['loc'])) + "']" if error['type'] == 'value_error.missing': print(f"Missing key at path \"{path_list_str}\"") elif error['type'].startswith('type_error'): expected_type = error['ctx']['expected_type'] print(f"Value at path \"{path_list_str}\" should be of type \"{expected_type}\" but got <class '{error['input'].__class__.__name__}'>")
Pydantic automatically handles nested validation, type conversion attempts, and detailed error messages. It's widely used in Python web frameworks and data pipelines for good reason.
内容的提问来源于stack exchange,提问作者TheAschr

