技术问询:更新数据时如何忽略但仍处理CSV中的非合规行
Since you're working with backfilled Google Finance data (given your parameters like EXCHANGE%3DINDEXSP, COLUMNS=DATE,OPEN) for a proof-of-concept visualization, dealing with messy or non-compliant CSV rows is super common. Here are a few actionable, code-friendly ways to skip bad rows while still processing the good ones:
1. Line-by-Line Validation with Manual Checks
Instead of feeding the entire CSV into a parser at once, read each line individually, validate it against your expected schema, and only process rows that pass. This gives you full control over what counts as "compliant."
For example, in Python:
import csv from datetime import datetime def is_valid_row(row): # Check if required columns exist (match your COLUMNS parameter) if "DATE" not in row or "OPEN" not in row: return False # Validate DATE is a parseable timestamp (adjust format if Google uses ISO instead of Unix) try: datetime.fromtimestamp(int(row["DATE"])) except (ValueError, TypeError): return False # Validate OPEN is a numeric value try: float(row["OPEN"]) except (ValueError, TypeError): return False return True with open("google_finance_data.csv", "r") as f: reader = csv.DictReader(f) valid_rows = [] for line_num, row in enumerate(reader, start=2): # Line 1 is header if is_valid_row(row): valid_rows.append(row) else: print(f"Skipping non-compliant row {line_num}: {row}") # Use valid_rows to update your visualization
2. Leverage CSV Parser Error Handling
Most CSV libraries let you catch parsing errors and skip problematic rows. For example, Python's built-in csv module raises csv.Error for malformed lines:
import csv with open("google_finance_data.csv", "r") as f: reader = csv.reader(f) header = next(reader) # Grab header first valid_rows = [header] for line_num, row in enumerate(reader, start=2): try: # Basic check: row has same column count as header if len(row) != len(header): raise csv.Error("Column count mismatch") valid_rows.append(row) except csv.Error as e: print(f"Error parsing row {line_num}: {e} — skipping") # Process valid_rows for your update
3. Post-Parsing Filtering (Quick & Dirty)
If you don't need strict upfront validation, parse the entire CSV first, then filter out rows that don't meet your criteria. This is simpler but uses more memory for large CSVs:
import pandas as pd # Ideal if you're already using pandas for visualization # Load CSV, replace obvious invalid values with NaN df = pd.read_csv("google_finance_data.csv", na_values=["", "N/A", "null"]) # Drop rows with missing critical data df = df.dropna(subset=["DATE", "OPEN"]) # Convert columns to correct types, forcing errors to NaN df["DATE"] = pd.to_datetime(df["DATE"], errors="coerce") df["OPEN"] = pd.to_numeric(df["OPEN"], errors="coerce") # Drop any remaining invalid rows df = df.dropna(subset=["DATE", "OPEN"]) # Use df to update your visualization
Key Notes for Your Use Case
- Since you're working with backfilled data, some rows might have missing values or malformed timestamps (especially around your
MARKET_OPEN_MINUTE=570andMARKET_CLOSE_MINUTE=960ranges). Add time-range checks to your validation if needed. - Since this is a proof-of-concept, don't over-engineer it—pick the approach that fits your existing workflow (e.g., pandas is perfect if you're already using it for visualization).
内容的提问来源于stack exchange,提问作者Arash Howaida

