大体积TSV转JSON的Python脚本报错问题求助
Hey there! Let’s tackle your TSV-to-JSON conversion issues head-on—those large files and finicky fields (date, time, time_taken, c_ip) can definitely throw a wrench in things, especially when you’re prepping data for JSON analysis testing. Here’s a breakdown of common fixes and answers to your likely questions:
Most errors with these fields stem from format mismatches between TSV (unstructured text) and JSON (strict typed data). Here’s how to fix each one:
date/time fields
TSV often stores dates/times in non-standard formats (e.g.,05-12-2024vs2024/12/05, or missing timezone info), which breaks JSON serialization.
Fix: Use a dedicated date-time library to standardize values, with error handling for invalid entries:from datetime import datetime def standardize_datetime(raw_date, raw_time): try: # Adjust format string to match your TSV's date/time structure combined = f"{raw_date} {raw_time}" return datetime.strptime(combined, "%d-%m-%Y %H:%M:%S").isoformat() except ValueError: # Log invalid entries or set to null for JSON compatibility return Nonetime_taken
This field often includes units (e.g.,150ms,2.5s) or non-numeric values, which JSON won’t accept as a number type.
Fix: Extract numeric values and convert to a float/int:def clean_time_taken(raw_time): if not raw_time: return None # Strip non-digit/decimal characters numeric_part = ''.join(c for c in raw_time if c.isdigit() or c == '.') return float(numeric_part) if numeric_part else Nonec_ip (client IP)
Issues usually come from extra data (e.g.,192.168.1.1:8080with port) or malformed IP strings.
Fix: Normalize to pure IP address, supporting both IPv4 and IPv6:def clean_ip(raw_ip): if not raw_ip: return None # Split off port if present return raw_ip.split(':')[0].strip()
With files up to 20GB, loading everything into memory will crash your script. Use these strategies:
- Process line-by-line instead of reading the entire file. Use Python’s built-in
csvmodule withDictReader(for TSV, setdelimiter='\t'):import csv import json def process_row(row): # Apply the cleaning functions above to each field return { "date_time": standardize_datetime(row["date"], row["time"]), "time_taken": clean_time_taken(row["time_taken"]), "c_ip": clean_ip(row["c_ip"]), # Add other fields as needed } with open("input.tsv", "r") as tsv_in, open("output.jsonl", "w") as json_out: reader = csv.DictReader(tsv_in, delimiter='\t') for row in reader: processed = process_row(row) json.dump(processed, json_out) json_out.write("\n") # Use JSON Lines format for easy incremental parsing - Use JSON Lines (.jsonl) instead of a single large JSON array. This format (one JSON object per line) is easier for analysis tools to handle incrementally, and avoids memory issues with huge arrays.
- Q: My script is way too slow for 20GB files. How can I speed it up?
A: Try using faster libraries likepandaswith chunking (setchunksize=10000to process in batches), or command-line tools likecsvkit+jqfor lightning-fast conversions. For Python, swappingDictReaderforcsv.reader(using column indices instead of names) can also boost speed. - Q: What if some rows are missing fields entirely?
A: Userow.get("field_name", None)instead of direct indexing to avoidKeyError, and set missing values tonullin JSON—most analysis tools handle nulls gracefully. - Q: Can I parallelize the conversion to save time?
A: Absolutely! Split large TSVs into smaller chunks first (use thesplitcommand on Unix-like systems:split -l 100000 input.tsv chunk_), then process each chunk in parallel with Python’smultiprocessingmodule. You can combine the .jsonl outputs later with a simplecatcommand.
内容的提问来源于stack exchange,提问作者Jason Belanger

