如何用Python正则从TXT文件生成JSON或嵌套字典列表文件
Hey there! I see you're dealing with a huge quasi-JSON text file and want to convert it into nested dictionaries/lists using Python—let's walk through practical solutions tailored to your needs, since you already know your way around pandas, numpy, and XML parsing.
1. Fix & Parse Quasi-JSON with Standard Libraries
First, if your file's only differences from valid JSON are small quirks (like single quotes instead of double, unquoted keys, or trailing commas), the easiest approach is to clean the text first, then use Python's built-in json module to parse it into a nested structure.
Here's a reusable cleaning function to handle common quasi-JSON issues:
import json import re def clean_quasi_json(text): # Replace single quotes with double quotes (standard JSON requires double quotes) cleaned = text.replace("'", '"') # Add double quotes around unquoted keys (e.g., `name:` → `"name":`) cleaned = re.sub(r'(\w+):', r'"\1":', cleaned) # Remove trailing commas before closing braces/brackets (e.g., `[1,2,]` → `[1,2]`) cleaned = re.sub(r',\s*([}\]])', r'\1', cleaned) return cleaned # For huge files: process line-by-line to avoid loading everything into memory with open('huge_quasi_json.txt', 'r') as f: for line in f: stripped_line = line.strip() if not stripped_line: continue try: cleaned_line = clean_quasi_json(stripped_line) nested_dict = json.loads(cleaned_line) # Do something with nested_dict here (e.g., analyze, store, etc.) print(nested_dict) except json.JSONDecodeError as e: print(f"Failed to parse line: {e}")
Note: This basic cleaner works for common cases. If your file has more unusual syntax (like comments, or multi-line strings), you might need to extend the regex rules.
2. Build Nested Structures Directly with Regular Expressions
If your quasi-JSON has a consistent structure but is too far from valid JSON to clean easily, you can use regex to capture key-value pairs and recursively build nested dictionaries/lists.
Here's a recursive parser example:
import re from typing import Dict, List, Any def parse_quasi_json_with_regex(text: str) -> Dict[str, Any]: # Regex patterns to match keys, values, dictionaries, and lists key_val_pattern = re.compile(r'(\w+|".+?"):\s*(.+?)(?=,\s*(\w+|".+?"):|\s*})', re.DOTALL) dict_pattern = re.compile(r'{(.+?)}', re.DOTALL) list_pattern = re.compile(r'\[(.+?)\]', re.DOTALL) def parse_value(value_str: str) -> Any: # Handle quoted strings if value_str.startswith(('"', "'")) and value_str.endswith(('"', "'")): return value_str.strip('"\'') # Handle numbers try: return int(value_str) except ValueError: try: return float(value_str) except ValueError: pass # Handle nested dictionaries if value_str.startswith('{'): dict_content = dict_pattern.search(value_str).group(1) return parse_quasi_json_with_regex(dict_content) # Handle nested lists if value_str.startswith('['): list_content = list_pattern.search(value_str).group(1) items = [parse_value(item.strip()) for item in re.split(r',\s*', list_content)] return items # Fallback for unrecognized values return value_str result = {} matches = key_val_pattern.findall(text) for key_match, value_match, _ in matches: key = key_match.strip('"') result[key] = parse_value(value_match.strip()) return result # Example usage sample_text = "{user: 'Bob', stats: {posts: 42, likes: 1234}, tags: ['python', 'data']}" parsed_data = parse_quasi_json_with_regex(sample_text) print(parsed_data)
Note: Regex works best for predictable, simple structures. If your quasi-JSON has deeply nested or ambiguous syntax, regex can get messy fast.
3. Handle Large Files Efficiently
Since you're working with an extremely large file, never load the entire file into memory at once. Stick to:
- Line-by-line processing (as shown in the first example)
- Chunked reading (if lines are too long, read in fixed-size chunks and process partial data carefully)
- Generators to yield parsed objects one at a time, instead of storing all of them in a list
4. For Complex Quasi-JSON: Use a Parsing Library
If your quasi-JSON has non-standard features (like comments, multi-line values, or custom syntax), regex and basic cleaning might not cut it. Instead, use a library like pyparsing to define custom grammar rules for your file format:
from pyparsing import ( Word, alphas, nums, Combine, Optional, Dict, Group, Forward, ZeroOrMore, QuotedString, White ) # Define grammar rules for your quasi-JSON number = Combine(Optional('-') + nums + Optional('.' + nums)) string = QuotedString("'") | QuotedString('"') key = Word(alphas, alphas+nums+'_') | QuotedString('"') value = Forward() dict_expr = Dict(ZeroOrMore(Group(key + Optional(White()) + ':' + Optional(White()) + value))) list_expr = Group('[' + ZeroOrMore(value + Optional(',' + White())) + ']') value << (string | number | dict_expr | list_expr) # Parse the text sample_text = "{user_name: 'Charlie', active: True, preferences: {theme: 'dark', notifications: [True, False]}}" parsed_data = dict_expr.parseString(sample_text).asDict() print(parsed_data)
This approach is far more robust for complex, non-standard formats than regex alone.
内容的提问来源于stack exchange,提问作者Python Spark

