如何高效将Python结构转换为serde_json::Value?(PyO3集成场景)
Hey there! Let's walk through all the practical ways to convert between Python native types and serde_json::Value for your PyO3-powered JSON Schema validator, plus the key tradeoffs to consider based on your use cases and existing benchmark data.
1. Direct JSON String Passing
This is the approach you already identified as the fastest option.
How it works
Have Python pass a pre-serialized JSON string directly to your Rust function, then parse it with serde_json::from_str to get a serde_json::Value.
Example code
use pyo3::prelude::*; use serde_json::Value; #[pyfunction] fn validate_json_string(py: Python, json_str: &str, schema: &str) -> PyResult<()> { let value: Value = serde_json::from_str(json_str) .map_err(|e| PyErr::new::<pyo3::exceptions::PyValueError, _>(format!("Invalid JSON: {}", e)))?; // Run your schema validation logic here Ok(()) }
Tradeoffs
- Pros:
- Unbeatable performance: Skips all Python-Rust type conversion overhead, relying solely on serde's optimized JSON parser.
- Minimal code complexity: No custom type mapping needed.
- Cons:
- Only useful if Python already has a JSON string. If you need to convert a Python native structure, calling
json.dumpsfirst kills performance (as your benchmarks showed) due to redundant serialization/deserialization. - Can't directly handle Python objects — they must be pre-serialized.
- Only useful if Python already has a JSON string. If you need to convert a Python native structure, calling
2. Implement Serialize for a PyAny Newtype
This is the approach you're considering, where you wrap a &PyAny and implement serde's Serialize trait to map Python types to serde_json::Value.
How it works
Create a newtype wrapper for &PyAny, then recursively serialize Python's built-in types to their serde_json equivalents (dict → Object, list/tuple → Array, etc.).
Example code
use pyo3::prelude::*; use serde::ser::{Serialize, Serializer, SerializeSeq, SerializeMap}; struct WrapPyAny<'a>(&'a PyAny); impl<'a> Serialize for WrapPyAny<'a> { fn serialize<S>(&self, serializer: S) -> Result<S::Ok, S::Error> where S: Serializer, { let py = self.0.py(); match self.0 { _ if self.0.is_none() => serializer.serialize_none(), _ if let Ok(b) = self.0.downcast::<PyBool>() => serializer.serialize_bool(b.is_true()), _ if let Ok(s) = self.0.downcast::<PyString>() => serializer.serialize_str(&s.to_string_lossy()), _ if let Ok(i) = self.0.downcast::<PyLong>() => { i.extract::<i64>() .map_err(|e| serde::ser::Error::custom(format!("Failed to extract int: {}", e))) .and_then(|val| serializer.serialize_i64(val)) }, _ if let Ok(f) = self.0.downcast::<PyFloat>() => { f.extract::<f64>() .map_err(|e| serde::ser::Error::custom(format!("Failed to extract float: {}", e))) .and_then(|val| serializer.serialize_f64(val)) }, _ if let Ok(list) = self.0.downcast::<PyList>() => { let mut seq = serializer.serialize_seq(Some(list.len()))?; for item in list.iter() { seq.serialize_element(&WrapPyAny(item))?; } seq.end() }, _ if let Ok(tuple) = self.0.downcast::<PyTuple>() => { let mut seq = serializer.serialize_seq(Some(tuple.len()))?; for item in tuple.iter() { seq.serialize_element(&WrapPyAny(item))?; } seq.end() }, _ if let Ok(dict) = self.0.downcast::<PyDict>() => { let mut map = serializer.serialize_map(Some(dict.len()))?; for (key, value) in dict.iter() { let key_str = key.downcast::<PyString>() .map_err(|e| serde::ser::Error::custom(format!("Dict key must be string: {}", e)))?; map.serialize_entry(&key_str.to_string_lossy(), &WrapPyAny(value))?; } map.end() }, _ => Err(serde::ser::Error::custom(format!( "Unsupported Python type: {}", self.0.get_type().name() ))), } } } // Usage in your validator #[pyfunction] fn validate_python_obj(py: Python, data: &PyAny, schema: &str) -> PyResult<()> { let value = serde_json::to_value(WrapPyAny(data)) .map_err(|e| PyErr::new::<pyo3::exceptions::PyValueError, _>(format!("Serialization failed: {}", e)))?; // Run validation Ok(()) }
Tradeoffs
- Pros:
- Great performance: Only slightly slower than direct JSON strings, and way faster than the
json.dumpsroundtrip. - Handles Python native structures directly, no extra work needed on the Python side.
- Full control over conversion logic (e.g., you can add support for
datetimeby serializing to ISO strings).
- Great performance: Only slightly slower than direct JSON strings, and way faster than the
- Cons:
- Requires manual maintenance of the serialization logic — easy to miss edge cases (like
bytesor custom Python types). - Recursive traversal of deep structures could lead to stack overflow (mitigate with iterative traversal if needed).
- Error handling adds boilerplate code.
- Requires manual maintenance of the serialization logic — easy to miss edge cases (like
3. Use the pyo3-serde Crate
This is a ready-made library that bridges PyO3 and serde, handling type conversions out of the box.
How it works
pyo3-serde provides from_py (Python → serde-compatible types) and to_py (serde-compatible types → Python) functions that handle all built-in types automatically.
Example code
use pyo3::prelude::*; use pyo3_serde::{from_py, to_py}; use serde_json::Value; #[pyfunction] fn validate_with_pyo3_serde(py: Python, data: &PyAny, schema: &str) -> PyResult<()> { let value: Value = from_py(data)?; // Run validation Ok(()) } // Reverse conversion for validation errors #[pyfunction] fn get_validation_errors(py: Python) -> PyResult<Py<PyAny>> { let error_value: Value = /* Your error value from validation */; let py_error_obj = to_py(py, &error_value)?; Ok(py_error_obj) }
Tradeoffs
- Pros:
- Zero custom conversion code — saves time and reduces maintenance overhead.
- Performance is on par with your custom
Serializeimplementation (it's optimized under the hood). - Supports bidirectional conversion, covering your need to turn validation errors back into Python objects.
- Extensible: You can add custom serde logic for non-standard types if needed.
- Cons:
- Adds an external dependency (make sure versions of
pyo3,serde_json, andpyo3-serdeare compatible). - Slightly less flexibility than a fully custom implementation (though this is rarely an issue for standard types).
- Adds an external dependency (make sure versions of
4. Convert via Python's json Module in Rust
This approach calls Python's json.dumps from Rust, then parses the resulting string into serde_json::Value.
How it works
Import Python's json module in your Rust function, call dumps on the input object, then parse the string.
Example code
use pyo3::prelude::*; use serde_json::Value; #[pyfunction] fn validate_via_python_json(py: Python, data: &PyAny, schema: &str) -> PyResult<()> { let json_module = PyModule::import(py, "json")?; let json_str = json_module.call_method1("dumps", (data,))?.extract::<String>()?; let value: Value = serde_json::from_str(&json_str)?; // Run validation Ok(()) }
Tradeoffs
- Pros:
- Reuses Python's mature JSON serialization logic, supporting any type that
json.dumpscan handle (including custom classes with__json__methods).
- Reuses Python's mature JSON serialization logic, supporting any type that
- Cons:
- Terrible performance: Double serialization (Python → string) and deserialization (string →
serde_json::Value) makes this the slowest option. It's completely unsuitable for your Hypothesis testing use case where speed is critical. - Relies on Python's GIL and adds overhead of cross-language function calls.
- Terrible performance: Double serialization (Python → string) and deserialization (string →
5. Manual PyO3 Type Mapping
Skip serde entirely and write a direct conversion function between PyAny and serde_json::Value.
How it works
Write a recursive function that matches Python types and constructs the corresponding serde_json::Value variant directly.
Example code snippet
use pyo3::prelude::*; use serde_json::Value; fn py_to_value(py: Python, obj: &PyAny) -> PyResult<Value> { if obj.is_none() { Ok(Value::Null) } else if let Ok(b) = obj.downcast::<PyBool>() { Ok(Value::Bool(b.is_true())) } else if let Ok(s) = obj.downcast::<PyString>() { Ok(Value::String(s.to_string_lossy().into_owned())) } else if let Ok(list) = obj.downcast::<PyList>() { let mut arr = Vec::new(); for item in list.iter() { arr.push(py_to_value(py, item)?); } Ok(Value::Array(arr)) } else if let Ok(dict) = obj.downcast::<PyDict>() { let mut map = serde_json::Map::new(); for (key, value) in dict.iter() { let key_str = key.downcast::<PyString>()?.to_string_lossy().into_owned(); map.insert(key_str, py_to_value(py, value)?); } Ok(Value::Object(map)) } // Handle other types (int, float, tuple) here else { Err(PyErr::new::<pyo3::exceptions::PyTypeError, _>(format!( "Unsupported type: {}", obj.get_type().name() ))) } }
Tradeoffs
- Pros:
- Maximum control over the conversion process — you can optimize hot paths (like large lists/dicts) for speed.
- No serde dependency overhead (though you're already using
serde_json, so this is minor).
- Cons:
- Most code-heavy option — you have to handle every type and edge case manually.
- Error handling is verbose, and it's easy to miss supported types.
- Requires writing a corresponding reverse function for converting
serde_json::Valueback to Python objects.
Web Request/Response Validation
- If input is a raw JSON string: Use Direct JSON String Passing — it's the fastest and simplest choice for request bodies.
- If input is a Python-native object: Use pyo3-serde for minimal code and solid performance, or your custom
Serializenewtype if you need precise control over how types are converted.
Hypothesis Property Testing
- Avoid
json.dumpsat all costs: It's too slow to generate a high volume of test cases. - If Hypothesis generates JSON strings: Stick with Direct JSON String Passing to maximize throughput.
- If Hypothesis generates Python structures: Use pyo3-serde or your custom
Serializenewtype — both will give you enough speed to outperform your current Python-based validator by a wide margin.
serde_json::Value → Python Objects For turning validation errors back into Python objects, your best options are:
pyo3-serde'sto_py: The simplest and most maintainable choice, directly mappingserde_json::Valueto Python's native types.- Manual mapping: Write a recursive function to convert each
Valuevariant to the corresponding Python type (good for full control, but requires more code). - Avoid JSON string roundtrips: Converting to a string and using Python's
json.loadsis slow and unnecessary — only use this if you have a legacy compatibility requirement.
内容的提问来源于stack exchange,提问作者Stranger6667

