如何用Python3提取静态站点文件中的JSON frontmatter及内容?
Great question! Unlike YAML/TOML frontmatter which has clear delimiters (--- or +++), JSON frontmatter doesn't have a built-in way to mark its end. The key here is to leverage JSON's strict syntax rules to find where the frontmatter ends and the content begins. Here are a couple of practical approaches to solve this:
Approach 1: Incremental Parsing (Simple & Reliable)
This method works by gradually reading more of the file from the start and trying to parse it as JSON until we succeed. It's straightforward and doesn't require any extra dependencies, making it perfect for most static site content files (which are usually not extremely large).
import json def parse_json_frontmatter(file_path): with open(file_path, 'r', encoding='utf-8') as f: full_content = f.read() # Try parsing incrementally from the start for end_index in range(1, len(full_content)): try: # Attempt to parse the substring as JSON frontmatter = json.loads(full_content[:end_index]) # If successful, extract the remaining content (trim whitespace) content_body = full_content[end_index:].strip() return frontmatter, content_body except json.JSONDecodeError: # Keep trying if parsing fails continue # Edge case: entire file is JSON try: return json.loads(full_content), "" except json.JSONDecodeError: raise ValueError("No valid JSON frontmatter found at the start of the file")
Pros & Cons
- Pros: Dead simple to implement, handles all valid JSON cases correctly (including nested structures, strings with special characters).
- Cons: Less efficient for very large files, since it tries parsing every possible substring length until it hits the right one.
Approach 2: Bracket Balance + Validation (Optimized)
For larger files, we can optimize by first finding the matching closing bracket/brace for the top-level JSON structure (either {} or []), then validating that substring as JSON. This cuts down on the number of parsing attempts we need to make.
We have to be careful to ignore brackets that are inside JSON strings though—here's a robust implementation:
import json def find_json_end_index(content): # Check if content starts with a valid top-level JSON structure if not content.startswith(('{', '[')): raise ValueError("File does not start with a valid JSON object or array") open_char = content[0] close_char = '}' if open_char == '{' else ']' balance = 1 in_string = False escape_active = False for idx, char in enumerate(content[1:], start=1): # Handle escaped characters (skip the next char if we see a backslash) if escape_active: escape_active = False continue # Toggle string state when we hit an unescaped quote if char == '"': in_string = not in_string # Mark escape state for the next character elif char == '\\': escape_active = True # Only update bracket balance if we're not inside a string elif not in_string: if char == open_char: balance += 1 elif char == close_char: balance -= 1 # When balance hits 0, we've found the end of the top-level JSON if balance == 0: return idx + 1 # Return index after the closing bracket # If we exit the loop without balancing, the JSON is invalid raise ValueError("Unclosed JSON structure—no matching closing bracket found") def parse_json_frontmatter(file_path): with open(file_path, 'r', encoding='utf-8') as f: full_content = f.read() try: json_end_idx = find_json_end_index(full_content) # Validate the extracted JSON substring frontmatter = json.loads(full_content[:json_end_idx]) # Extract and trim the remaining content content_body = full_content[json_end_idx:].strip() return frontmatter, content_body except (json.JSONDecodeError, ValueError) as e: raise ValueError(f"Failed to parse JSON frontmatter: {str(e)}")
Pros & Cons
- Pros: Much faster for large files, since it directly locates the likely end of the JSON instead of brute-forcing every substring.
- Cons: Slightly more complex code, but still manageable.
Key Notes
- Make sure the JSON frontmatter is at the very start of the file—no leading spaces, newlines, or other characters, otherwise parsing will fail.
- Standard JSON doesn't support comments. If your files have JSON with comments, you'll need to pre-process the content to strip comments before parsing (though this isn't recommended, as it deviates from JSON specs).
- Always wrap parsing in try/except blocks to handle invalid JSON gracefully and give clear error messages to users.
内容的提问来源于stack exchange,提问作者user9703562

