You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则从TXT文件生成JSON或嵌套字典列表文件

Hey there! I see you're dealing with a huge quasi-JSON text file and want to convert it into nested dictionaries/lists using Python—let's walk through practical solutions tailored to your needs, since you already know your way around pandas, numpy, and XML parsing.

1. Fix & Parse Quasi-JSON with Standard Libraries

First, if your file's only differences from valid JSON are small quirks (like single quotes instead of double, unquoted keys, or trailing commas), the easiest approach is to clean the text first, then use Python's built-in json module to parse it into a nested structure.

Here's a reusable cleaning function to handle common quasi-JSON issues:

import json
import re

def clean_quasi_json(text):
    # Replace single quotes with double quotes (standard JSON requires double quotes)
    cleaned = text.replace("'", '"')
    # Add double quotes around unquoted keys (e.g., `name:` → `"name":`)
    cleaned = re.sub(r'(\w+):', r'"\1":', cleaned)
    # Remove trailing commas before closing braces/brackets (e.g., `[1,2,]` → `[1,2]`)
    cleaned = re.sub(r',\s*([}\]])', r'\1', cleaned)
    return cleaned

# For huge files: process line-by-line to avoid loading everything into memory
with open('huge_quasi_json.txt', 'r') as f:
    for line in f:
        stripped_line = line.strip()
        if not stripped_line:
            continue
        try:
            cleaned_line = clean_quasi_json(stripped_line)
            nested_dict = json.loads(cleaned_line)
            # Do something with nested_dict here (e.g., analyze, store, etc.)
            print(nested_dict)
        except json.JSONDecodeError as e:
            print(f"Failed to parse line: {e}")

Note: This basic cleaner works for common cases. If your file has more unusual syntax (like comments, or multi-line strings), you might need to extend the regex rules.

2. Build Nested Structures Directly with Regular Expressions

If your quasi-JSON has a consistent structure but is too far from valid JSON to clean easily, you can use regex to capture key-value pairs and recursively build nested dictionaries/lists.

Here's a recursive parser example:

import re
from typing import Dict, List, Any

def parse_quasi_json_with_regex(text: str) -> Dict[str, Any]:
    # Regex patterns to match keys, values, dictionaries, and lists
    key_val_pattern = re.compile(r'(\w+|".+?"):\s*(.+?)(?=,\s*(\w+|".+?"):|\s*})', re.DOTALL)
    dict_pattern = re.compile(r'{(.+?)}', re.DOTALL)
    list_pattern = re.compile(r'\[(.+?)\]', re.DOTALL)

    def parse_value(value_str: str) -> Any:
        # Handle quoted strings
        if value_str.startswith(('"', "'")) and value_str.endswith(('"', "'")):
            return value_str.strip('"\'')
        # Handle numbers
        try:
            return int(value_str)
        except ValueError:
            try:
                return float(value_str)
            except ValueError:
                pass
        # Handle nested dictionaries
        if value_str.startswith('{'):
            dict_content = dict_pattern.search(value_str).group(1)
            return parse_quasi_json_with_regex(dict_content)
        # Handle nested lists
        if value_str.startswith('['):
            list_content = list_pattern.search(value_str).group(1)
            items = [parse_value(item.strip()) for item in re.split(r',\s*', list_content)]
            return items
        # Fallback for unrecognized values
        return value_str

    result = {}
    matches = key_val_pattern.findall(text)
    for key_match, value_match, _ in matches:
        key = key_match.strip('"')
        result[key] = parse_value(value_match.strip())
    return result

# Example usage
sample_text = "{user: 'Bob', stats: {posts: 42, likes: 1234}, tags: ['python', 'data']}"
parsed_data = parse_quasi_json_with_regex(sample_text)
print(parsed_data)

Note: Regex works best for predictable, simple structures. If your quasi-JSON has deeply nested or ambiguous syntax, regex can get messy fast.

3. Handle Large Files Efficiently

Since you're working with an extremely large file, never load the entire file into memory at once. Stick to:

  • Line-by-line processing (as shown in the first example)
  • Chunked reading (if lines are too long, read in fixed-size chunks and process partial data carefully)
  • Generators to yield parsed objects one at a time, instead of storing all of them in a list

4. For Complex Quasi-JSON: Use a Parsing Library

If your quasi-JSON has non-standard features (like comments, multi-line values, or custom syntax), regex and basic cleaning might not cut it. Instead, use a library like pyparsing to define custom grammar rules for your file format:

from pyparsing import (
    Word, alphas, nums, Combine, Optional, Dict, Group,
    Forward, ZeroOrMore, QuotedString, White
)

# Define grammar rules for your quasi-JSON
number = Combine(Optional('-') + nums + Optional('.' + nums))
string = QuotedString("'") | QuotedString('"')
key = Word(alphas, alphas+nums+'_') | QuotedString('"')
value = Forward()
dict_expr = Dict(ZeroOrMore(Group(key + Optional(White()) + ':' + Optional(White()) + value)))
list_expr = Group('[' + ZeroOrMore(value + Optional(',' + White())) + ']')
value << (string | number | dict_expr | list_expr)

# Parse the text
sample_text = "{user_name: 'Charlie', active: True, preferences: {theme: 'dark', notifications: [True, False]}}"
parsed_data = dict_expr.parseString(sample_text).asDict()
print(parsed_data)

This approach is far more robust for complex, non-standard formats than regex alone.


内容的提问来源于stack exchange,提问作者Python Spark

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:13:36