Python csv.reader需求:返回带引号移除标记的单元格元组
Great question! Let's figure out how to tweak csv.reader to return not just the parsed cell value, but also a boolean flag indicating whether quotes were stripped from the original cell.
Approach: Compare Raw vs. Parsed Cells
The core trick is to compare the raw, unprocessed cell string (including any quotes) with the final parsed value from csv.reader. If the parsed value matches the raw string with outer quotes removed and internal escaped quotes fixed, we know quotes were stripped.
Here's a practical implementation:
import csv from io import StringIO def quoted_flag_reader(source, delimiter='\t'): # Handle both file objects and strings (for non-seekable streams like stdin) if isinstance(source, str): source = StringIO(source) # First reader: gets the normal parsed values (quotes stripped, escapes handled) normal_reader = csv.reader(source, delimiter=delimiter) # Reset the stream to read raw, unprocessed content source.seek(0) # Second reader: gets raw cell strings (no quote processing at all) raw_reader = csv.reader( source, delimiter=delimiter, quoting=csv.QUOTE_NONE, quotechar='', escapechar='' ) # Iterate through both readers row by row for normal_row, raw_row in zip(normal_reader, raw_reader): flagged_row = [] for normal_val, raw_val in zip(normal_row, raw_row): was_quoted = False # Check if the raw cell is properly wrapped in matching quotes if len(raw_val) >= 2 and raw_val.startswith('"') and raw_val.endswith('"'): # Replicate how csv.reader processes quoted cells internally processed_raw = raw_val[1:-1].replace('""', '"') # If this matches the parsed value, quotes were stripped during parsing if processed_raw == normal_val: was_quoted = True flagged_row.append( (normal_val, was_quoted) ) yield flagged_row
Test It With Your Example
Suppose your tab-separated file has this line:
"123" 123 """123"""
Run this test code:
# Test with a local file with open('test.csv', 'r', newline='') as f: for row in quoted_flag_reader(f, delimiter='\t'): print(row) # Or test directly with a string test_line = '"123"\t123\t"""123"""' for row in quoted_flag_reader(test_line, delimiter='\t'): print(row)
Expected Output
[('123', True), ('123', False), ('"123"', True)]
Let's break down the results:
('123', True): Original cell was"123"—csv.readerstripped the outer quotes, so the flag isTrue.('123', False): Original cell was123(no quotes), so no stripping happened.('"123"', True): Original cell was"""123"""—csv.readerstripped the outer quotes and replaced the inner""with", resulting in"123", so the flag isTrue.
Key Notes
- Non-seekable Streams: If you're reading from a stream that can't be reset (like
sys.stdin), pass the entire content as a string to the function (we added support for that withStringIO). - Dialect Adjustments: If your CSV uses commas instead of tabs, or a different quote character, adjust the
delimiterandquotecharparameters accordingly. - Malformed Quotes: The function ignores cells with mismatched quotes (e.g.,
"123without a closing quote) becausecsv.readerdoesn't strip quotes from invalidly quoted content anyway.
内容的提问来源于stack exchange,提问作者user2567875
相关产品推荐
相关产品推荐

