如何使用Python解析并清理含JSON原始数据的文本文件?
Alright, let's break down how to parse and clean this escaped JSON data you've got. It looks like you're dealing with a string full of escape sequences (\x22 for double quotes, \x0A for newlines) that also got truncated at the end. Here's a practical, step-by-step approach to fix and process it:
Step 1: Decode Escape Sequences
First, we need to convert those escape codes into actual readable characters. Most programming languages have built-in methods for this—here's how to do it in Python, which is great for quick data manipulation:
# Your raw escaped data raw_data = '{\x0A \x22identifier\x22: {\x0A \x22company_code\x22: \x22TSC\x22,\x0A \x22product_type\x22: \x22airtime-ctg\x22,\x0A \x22host_type\x22: \x22android\x22\x0A },\x0A \x22id\x22: {\x0A \x22type\x22: \x22guest\x22,\x0A \x22group\x22: \x22guest\x22,\x0A \x22uuid\x22: \x221a0d4d6e-0c00-11e7-a16f-0242ac110002\x22,\x0A \x22device_id\x22: \x22423e49efa4b8b013\x2...' # Decode escape sequences to get a human-readable JSON structure decoded_data = raw_data.encode('utf-8').decode('unicode-escape') print(decoded_data)
Running this will give you a partially formatted JSON string like this:
{ "identifier": { "company_code": "TSC", "product_type": "airtime-ctg", "host_type": "android" }, "id": { "type": "guest", "group": "guest", "uuid": "1a0d4d6e-0c00-11e7-a16f-0242ac110002", "device_id": "423e49efa4b8b013...
Step 2: Fix the Truncated JSON
Notice the data cuts off mid-string (\x2...)? That means the JSON is incomplete and won't parse correctly. You have two options here:
- Manually repair it: If you know the original complete data, fill in the missing parts (close the
device_idstring, theidobject, and the top-level JSON object). A fixed version might look like this:{ "identifier": { "company_code": "TSC", "product_type": "airtime-ctg", "host_type": "android" }, "id": { "type": "guest", "group": "guest", "uuid": "1a0d4d6e-0c00-11e7-a16f-0242ac110002", "device_id": "423e49efa4b8b013" } } - Use an auto-repair tool: If you don't have the full original data, libraries like
jsonrepaircan automatically fix common issues like truncated strings or missing brackets:from jsonrepair import repair_json fixed_json = repair_json(decoded_data) print(fixed_json)
Step 3: Parse and Clean the Data
Once you have a valid, complete JSON string, you can parse it into a structured object (like a Python dictionary) and clean it to fit your needs. For example, you might want to flatten nested fields, remove unused data, or standardize field names:
import json # Use the fixed JSON string from Step 2 fixed_json_str = '''{ "identifier": { "company_code": "TSC", "product_type": "airtime-ctg", "host_type": "android" }, "id": { "type": "guest", "group": "guest", "uuid": "1a0d4d6e-0c00-11e7-a16f-0242ac110002", "device_id": "423e49efa4b8b013" } }''' # Parse JSON into a dictionary data_dict = json.loads(fixed_json_str) # Clean and restructure the data (customize this based on your needs) cleaned_data = { 'company_code': data_dict['identifier']['company_code'], 'product_type': data_dict['identifier']['product_type'], 'device_type': data_dict['identifier']['host_type'], 'user_type': data_dict['id']['type'], 'user_uuid': data_dict['id']['uuid'] } print(cleaned_data)
This will output a clean, flat structure ready for analysis or storage:
{ 'company_code': 'TSC', 'product_type': 'airtime-ctg', 'device_type': 'android', 'user_type': 'guest', 'user_uuid': '1a0d4d6e-0c00-11e7-a16f-0242ac110002' }
Quick Notes
- If your data is stored in a text file, read it directly into your script first instead of copying the escaped string manually.
- Adjust the cleaning logic to match your specific use case—you might need to filter invalid values, convert data types, or keep additional fields.
内容的提问来源于stack exchange,提问作者tintin

