Python2与Python3脚本间的数据编列方案咨询
Hey there! Let's work through how to handle data marshaling between your Python 2 script and the Python 3 subprocess you're calling. Since you're already using subprocess to kick off the Python 3 code, here are the most practical, reliable approaches to pass data back and forth:
1. JSON (The Go-To for Most Scenarios)
JSON is hands down the best choice here—it’s cross-version compatible, human-readable, and supported natively by both Python 2 and 3. It works for all common data types (strings, numbers, lists, dictionaries, booleans) and avoids the compatibility headaches of Python-specific formats.
Python 2 Script (Sender/Receiver)
import json import subprocess # Prepare your data to send data_to_transfer = { "user_id": 123, "preferences": ["dark_mode", "notifications"], "is_active": True } # Serialize to JSON string json_payload = json.dumps(data_to_transfer) # Call Python 3 script, pass data via stdin, capture stdout/stderr proc = subprocess.Popen( ["/path/to/python3", "python3_script.py"], stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE ) stdout, stderr = proc.communicate(input=json_payload.encode("utf-8")) # Handle the response from Python 3 if proc.returncode == 0: response_data = json.loads(stdout.decode("utf-8")) print("Got back from Python 3:", response_data) else: print(f"Error running Python 3 script: {stderr.decode('utf-8')}")
Python 3 Script (Receiver/Sender)
import json import sys # Read JSON data from stdin input_data = json.load(sys.stdin) print("Received from Python 2:", input_data) # Process the data (example: modify preferences) input_data["preferences"].append("auto_save") result = {"status": "success", "updated_data": input_data} # Serialize result and send back via stdout json_result = json.dumps(result) sys.stdout.write(json_result)
Pro tips:
- Always use UTF-8 encoding to avoid character issues across versions.
- For non-JSON-native types (like
datetimeobjects), convert them to ISO strings before serialization, then parse them back on the other end.
2. Pickle (Python-Specific, Use With Care)
If you need to transfer Python-specific types that JSON can’t handle (like custom objects), pickle is an option—but it has major compatibility caveats between Python 2 and 3. To make it work:
- Use
protocol=2(the highest protocol compatible with both versions) when pickling. - On Python 3, specify
encoding="latin1"when unpickling to handle Python 2 strings correctly.
Python 2 Side
import pickle import subprocess data = {"custom_obj": MyCustomClass(), "values": (1, 2, 3)} # Pickle with compatible protocol pickled_data = pickle.dumps(data, protocol=2) proc = subprocess.Popen( ["/path/to/python3", "python3_script.py"], stdin=subprocess.PIPE, stdout=subprocess.PIPE ) stdout, _ = proc.communicate(input=pickled_data) # Unpickle the response result = pickle.loads(stdout)
Python 3 Side
import pickle import sys # Load pickled data with correct encoding input_data = pickle.load(sys.stdin, encoding="latin1") print("Received data:", input_data) # Process and send back processed = {"result": input_data["values"] + (4, 5)} sys.stdout.write(pickle.dumps(processed, protocol=2))
Warning: Never unpickle data from untrusted sources—it’s a security risk, as malicious pickle data can execute arbitrary code.
3. Command-Line Arguments (For Tiny, Simple Data)
If you’re only passing small, simple values (like single strings or numbers), you can send them directly as command-line arguments. Just be mindful of escaping special characters (like spaces or quotes).
Python 2 Script
import subprocess # Pass simple arguments username = "jane_doe" user_age = "30" subprocess.call(["/path/to/python3", "python3_script.py", username, user_age])
Python 3 Script
import sys # Access arguments via sys.argv username = sys.argv[1] user_age = int(sys.argv[2]) print(f"Hello {username}, you are {user_age} years old!")
Limitations: This isn’t feasible for large or complex data structures—stick to this only for trivial data.
4. Temporary Files (For Large Datasets)
If you’re working with massive data (like large CSV datasets, image blobs, or huge lists), passing data via stdin might hit memory limits. Instead, write the data to a temporary file, pass the file path to the Python 3 script, and have it write results to another file.
Python 2 Script
import subprocess import tempfile import json # Create a large dataset large_data = {"dataset": [i for i in range(1000000)]} # Write to a temporary file with tempfile.NamedTemporaryFile(mode='w', delete=False) as temp_file: json.dump(large_data, temp_file) # Call Python 3 script with the temp file path subprocess.call(["/path/to/python3", "python3_script.py", temp_file.name]) # Read the result file with open("processed_result.json", 'r') as result_file: result = json.load(result_file)
Python 3 Script
import json import sys # Read data from the temp file input_file_path = sys.argv[1] with open(input_file_path, 'r') as input_file: data = json.load(input_file) # Process the large dataset (example: calculate sum) processed_sum = sum(data["dataset"]) # Write result to a file with open("processed_result.json", 'w') as result_file: json.dump({"total_sum": processed_sum}, result_file)
Note: Remember to clean up temporary files after use (the delete=False flag keeps the file around until you manually delete it).
To recap: JSON is your best bet for most cases due to its compatibility, safety, and readability. Use pickle only if you absolutely need Python-specific types, and avoid it with untrusted data. For tiny data, command-line arguments work, and for huge datasets, temporary files are the way to go.
内容的提问来源于stack exchange,提问作者Roel

