如何将含唯一标识符的JSON批量解析为表格格式?
If you need to parse a folder of JSON files (each containing top-level unique block IDs) into a structured table, here's a straightforward Python solution that extracts key fields and outputs a Markdown table (or CSV if preferred).
Step 1: Understand the JSON Structure
From your example, each JSON file has top-level keys as unique block IDs, with each value containing:
category: Type of block (e.g.,chapter)children: List of child block IDsmetadata: Includesdisplay_nameandstarttimestamp
Step 2: Python Script to Process Files & Generate Table
This script will iterate over all JSON files in your target folder, extract relevant data, and generate a Markdown table. No external libraries are required (though pandas can simplify things if you need more advanced table handling).
import os import json def process_json_files(folder_path): """Process all JSON files in the folder and collect table data.""" table_rows = [] for filename in os.listdir(folder_path): if not filename.endswith(".json"): continue file_path = os.path.join(folder_path, filename) with open(file_path, "r", encoding="utf-8") as f: try: json_data = json.load(f) except json.JSONDecodeError: print(f"Skipping invalid JSON file: {filename}") continue # Iterate over each top-level block in the JSON for block_id, block_details in json_data.items(): # Extract fields, handle missing values gracefully row = { "Block ID": block_id, "Category": block_details.get("category", "N/A"), "Display Name": block_details["metadata"].get("display_name", "N/A"), "Start Date": block_details["metadata"].get("start", "N/A"), "Children": ", ".join(block_details.get("children", [])) or "N/A" } table_rows.append(row) return table_rows def generate_markdown_table(data): """Convert collected data into a Markdown table string.""" if not data: return "No valid data found in JSON files." # Get column headers from the first row headers = list(data[0].keys()) # Build Markdown table markdown_table = "| " + " | ".join(headers) + " |\n" markdown_table += "| " + " | ".join(["---"] * len(headers)) + " |\n" for row in data: # Convert each value to string to handle missing data row_values = [str(row[header]) for header in headers] markdown_table += "| " + " | ".join(row_values) + " |\n" return markdown_table # ---------------------- # Usage Example # ---------------------- if __name__ == "__main__": # Replace with your actual folder path target_folder = "/path/to/your/json/files" # Process files and generate table table_data = process_json_files(target_folder) markdown_output = generate_markdown_table(table_data) # Print the table to console print(markdown_output) # Or save to a Markdown file with open("converted_table.md", "w", encoding="utf-8") as f: f.write(markdown_output)
Step 3: Example Output
Based on your sample JSON snippet, the generated Markdown table would look like this:
| Block ID | Category | Display Name | Start Date | Children |
|---|---|---|---|---|
| block-v1:SampleData-type@chapter+block@14a0423ddf4a4d90926fb348e86a6232 | chapter | XYZZ | 2017-02-13T07:00:00Z | block-v1:SampleData-type@sequential+block@0fd2ac771bd141f384b8a3c628207d1d |
Notes
- The script handles invalid JSON files by skipping them and printing a warning.
- Missing fields are replaced with
N/Ato keep the table consistent. - If you prefer a CSV output instead of Markdown, you can use Python's
csvmodule to write thetable_rowslist directly to a .csv file.
内容的提问来源于stack exchange,提问作者pooja kosala

