如何在Python中将持续写入的无括号JSON文本文件转为JSON数组?
Alright, I’ve got you covered here. Let’s break down how to handle this tricky scenario—since you can’t modify the file being written by another program, and you need to grab the latest JSON entries as they come in, here’s a practical, incremental approach:
The key here is to read the file incrementally instead of trying to parse the whole thing as a single JSON array. We’ll track our position in the file so we only read new content each time, then parse each individual JSON object as it’s added—no need to add those missing brackets to the original file.
Step 1: Track Your Read Position
First, we need to remember where we left off the last time we read the file. This way, every time we check for updates, we don’t re-read the entire file—just the new stuff that’s been added since our last check.
Step 2: Parse New Chunks Safely
The file writes JSON objects separated by commas, so we’ll collect the new content, clean up any partial entries (in case the other program is still writing), then parse each valid JSON object.
Example Implementation
Here’s a Python script that does exactly this, with comments explaining each part:
import json import time def follow_json_updates(file_path, start_position=0): # Track how far we've read into the file last_read_pos = start_position while True: with open(file_path, 'r') as f: # Jump to the last position we finished reading f.seek(last_read_pos) # Grab all new content added since then new_content = f.read() # If there's no new content, wait a second and check again if not new_content: time.sleep(1) continue # Update our position to the current end of the file last_read_pos = f.tell() # Clean up trailing commas (in case the file ends mid-write) cleaned_content = new_content.rstrip(', ') # Split into individual JSON object strings entry_strings = cleaned_content.split(', ') for entry_str in entry_strings: try: # Parse the JSON string into a Python dict entry = json.loads(entry_str) # Do whatever you need with the latest entry here print("New entry received:", entry) except json.JSONDecodeError: # If the entry is incomplete (still being written), skip it for now # We'll catch it on the next loop when it's fully saved pass # Allow exiting with Ctrl+C try: time.sleep(1) except KeyboardInterrupt: print("\nStopping update checker.") break # Return the last position so you can resume later if needed return last_read_pos # How to use it: # First run: start reading from the beginning of the file # final_position = follow_json_updates("your_data_file.txt") # If you restart the script later, pass the final_position to pick up where you left off # final_position = follow_json_updates("your_data_file.txt", start_position=final_position)
Key Details & Edge Cases
- Handling Partial Entries: If the other program is still writing an entry (so the JSON is cut off mid-write),
json.loads()will throw an error. We just skip that entry and pick it up in the next loop iteration once it’s fully saved. - Persistent Position Tracking: If you want to resume reading after restarting your script, save the
last_read_posvalue to a small text file (e.g.,open("last_position.txt", "w").write(str(last_read_pos))), then read it back when starting. - Complex JSON Objects: If your JSON objects might contain commas inside strings (like
{"note": "Bought apples, oranges"}), the simplesplit(', ')will break. For that, usejson.JSONDecoderto incrementally parse the content—here’s a quick tweak for that scenario:decoder = json.JSONDecoder() pos = 0 while pos < len(cleaned_content): try: entry, pos = decoder.raw_decode(cleaned_content, pos) print("New entry received:", entry) # Skip the comma separator if present if pos < len(cleaned_content) and cleaned_content[pos] == ',': pos += 2 # Skip ', ' except json.JSONDecodeError: break - Performance: The
time.sleep(1)prevents your script from spamming the file system with constant checks. Adjust this based on how often new entries are added (e.g., use0.5if entries come in fast, or5if they’re slow).
内容的提问来源于stack exchange,提问作者Jith

