Python 3读取含NUL的CSV文件时遇_csv.Error: line contains NULL错误求助
Hey Avishay, I’ve run into this exact NUL character issue with CSV files before—those pesky \x00 bytes break the standard csv module because they’re not allowed in valid CSV. Let’s get your code working again with a few straightforward fixes.
The Root Problem
CSV doesn’t permit NUL (null) characters, so when csv.DictReader encounters them, it’ll throw an error like _csv.Error: line contains NUL. Your current code doesn’t handle these invalid bytes, which is why you’re hitting roadblocks.
Solution 1: Filter NUL Characters On-the-Fly (Best for Large Files)
Instead of loading the entire file into memory, use a generator to strip NUL characters as you read each line. This is efficient for big logs:
import csv with open(log_path, 'r', encoding='utf-8') as csv_file: # Strip NUL characters from each line as we read it cleaned_lines = (line.replace('\x00', '') for line in csv_file) log_reader = csv.DictReader(cleaned_lines) for line in log_reader: # Quick note: Your original code had `== str` which compares to the string type # Make sure you compare to an actual string value, like 'your_target_string' if line['Addition Information'] == 'your_target_string': # Do your processing here print(line)
Solution 2: Clean the Entire File at Once (Good for Small Files)
If your log file isn’t too large, you can read the whole thing, strip NULs, and feed it to csv.DictReader using StringIO:
import csv from io import StringIO with open(log_path, 'r', encoding='utf-8') as f: # Remove all NUL characters from the entire file content cleaned_content = f.read().replace('\x00', '') # Treat the cleaned string as a file object log_reader = csv.DictReader(StringIO(cleaned_content)) for line in log_reader: if line['Addition Information'] == 'your_target_string': # Process the line pass
Quick Fix for Your Original Code’s Typo
Don’t forget: your original condition line['Addition Information'] == str compares the field value to the Python str type, not a string literal. That will never evaluate to True—make sure you replace str with the actual string you’re checking for (e.g., 'error' or 'warning').
Why Ditch codecs.open?
Python 3’s built-in open() supports the encoding parameter directly, so it’s simpler and avoids any unexpected encoding handling quirks that codecs.open might introduce. If your file uses a non-UTF-8 encoding (like UTF-16), just adjust the encoding argument (e.g., encoding='utf-16').
内容的提问来源于stack exchange,提问作者Avishay Cohen

