如何从JSON文件提取值?正则是否可行?超大规模无序文件解析求助
Great question—let’s break this down step by step since you’re dealing with a massive, messy dataset and a couple of related tasks.
First, let’s tackle the 20M+ line file. The key here is working with the data in streams (to avoid memory overload) and tailoring your extraction to each field’s predictability:
Start with a sample analysis
Don’t dive into the full file first. Grab 100–1000 random lines to map out patterns:- Standard fields like emails, IPs, and phones likely follow consistent formats.
- Names and addresses might appear with prefixes (e.g.,
Name: Jane Smith,Zip: 90210) or as free-form text—note these variations.
Use stream-based tools
Loading the entire file into memory will crash most systems. Instead, process line-by-line:- For Python, use
fileinputor plain line iteration withopen(). - For command-line workflows,
awkorgrepcan handle simple matches, but Python is more flexible for complex extraction.
- For Python, use
Field-specific extraction strategies
- Standard fields (Email, IP, Phone): Use regular expressions—they’re perfect here. Precompile your regexes to speed up processing:
import re import fileinput # Precompile regex patterns for speed email_re = re.compile(r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}') ipv4_re = re.compile(r'\b(?:\d{1,3}\.){3}\d{1,3}\b') # Adjust phone regex to match your region's format (example for US numbers) phone_re = re.compile(r'(\+1\s?)?(?:\(\d{3}\)|\d{3})[-.\s]?\d{3}[-.\s]?\d{4}') with open('extracted_data.csv', 'w') as outfile: outfile.write("Email,IP,Phone,FirstName,LastName,Zip,Address1,Address2,City\n") for line in fileinput.input('your_large_file.txt'): line = line.strip() if not line: continue # Extract standard fields email = email_re.search(line).group() if email_re.search(line) else '' ip = ipv4_re.search(line).group() if ipv4_re.search(line) else '' phone = phone_re.search(line).group() if phone_re.search(line) else '' # For names/addresses: # If they have clear prefixes (e.g., "FirstName: John"), use regex to capture values # If free-form, use a lightweight NER tool like spaCy in stream mode (avoid loading the whole file) first_name = '' last_name = '' zip_code = '' addr1 = '' addr2 = '' city = '' # Example: Extract name if formatted as "Name: John Doe" name_match = re.search(r'Name: (\w+) (\w+)', line) if name_match: first_name, last_name = name_match.groups() # Write results (handle missing values as empty strings) outfile.write(f"{email},{ip},{phone},{first_name},{last_name},{zip_code},{addr1},{addr2},{city}\n") - Names/Addresses: If they’re free-form, regex will only get you so far. For better accuracy, use a named entity recognition (NER) library like spaCy—process lines one at a time to avoid memory bloat.
- Standard fields (Email, IP, Phone): Use regular expressions—they’re perfect here. Precompile your regexes to speed up processing:
JSON extraction depends on how your file is structured:
- JSON Lines (one object per line): This is ideal for large files. Use Python’s built-in
jsonmodule to parse line-by-line:import json with open('data.jsonl', 'r') as infile, open('json_output.csv', 'w') as outfile: outfile.write("Email,IP,Name\n") for line in infile: try: obj = json.loads(line) email = obj.get('email', '') ip = obj.get('ip', '') name = obj.get('name', '') outfile.write(f"{email},{ip},{name}\n") except json.JSONDecodeError: print(f"Skipping invalid JSON line: {line[:50]}...") - Single large JSON array: Use a streaming JSON parser like
ijsonto avoid loading the entire array into memory:import ijson with open('large.json', 'r') as infile, open('json_output.csv', 'w') as outfile: outfile.write("Email,IP,Name\n") # Parse each item in the top-level array for obj in ijson.items(infile, 'item'): email = obj.get('email', '') ip = obj.get('ip', '') name = obj.get('name', '') outfile.write(f"{email},{ip},{name}\n") - Command-line shortcut: Use
jqfor quick extractions. For example, to output emails and IPs as CSV:jq -r '.[] | [.email, .ip] | @csv' large.json > output.csv
Short answer: Yes, but with caveats:
- Perfect for standard fields: Emails, IPs, and phones have well-defined formats—regex is the most efficient way to extract these.
- Limited for non-standard fields: Names and addresses can vary wildly (e.g., hyphenated last names, street abbreviations). Regex works only if you can define clear patterns (like prefixes or consistent formatting). For free-form text, combine regex with NER or rule-based logic for better results.
- Pro tip: Always precompile regex patterns in Python—this drastically speeds up processing for large datasets.
内容的提问来源于stack exchange,提问作者RayCrush

