使用TextIOWrapper的csv.reader为何包含换行符?如何处理?
解决方案
不用手动遍历所有行替换,你可以通过修改ZIP处理函数中读取CSV内容的方式,模仿单个CSV函数的iter_lines()行为,自动处理多余换行符。
方法1:生成器逐行处理(适合大文件)
直接对ZIP内的CSV字节流逐行解码并清理换行符,避免一次性加载全部内容占用内存:
@request_exceptions def download_and_process_zip(client, url): with closing(client.get(url, stream=True)) as response: response.raise_for_status() with io.BytesIO(response.content) as buffer: with zipfile.ZipFile(buffer, 'r') as z: for filename in z.namelist(): base_filename, file_extension = os.path.splitext(filename) model_class = apps.get_model(base_filename) if file_extension == '.csv': with z.open(filename) as csv_file: # 逐行解码,替换连续换行并去除行尾换行符 cleaned_lines = ( line.decode('utf-8').replace('\n\n', '').rstrip('\n') for line in csv_file ) reader = csv.reader(cleaned_lines) process_copy_from_csv(model_class, reader)
方法2:批量读取处理(适合小文件)
如果CSV文件体积较小,可以一次性读取全部内容,批量替换多余换行后再拆分成行:
@request_exceptions def download_and_process_zip(client, url): with closing(client.get(url, stream=True)) as response: response.raise_for_status() with io.BytesIO(response.content) as buffer: with zipfile.ZipFile(buffer, 'r') as z: for filename in z.namelist(): base_filename, file_extension = os.path.splitext(filename) model_class = apps.get_model(base_filename) if file_extension == '.csv': with z.open(filename) as csv_file: text_wrapper = io.TextIOWrapper(csv_file, encoding='utf-8') # 读取全部内容并替换连续换行 cleaned_content = text_wrapper.read().replace('\n\n', '') lines = cleaned_content.splitlines() reader = csv.reader(lines) process_copy_from_csv(model_class, reader)
补充说明
- 若需要处理
\r\n\r\n这类跨平台连续换行,可导入re模块,将替换逻辑改为re.sub(r'(\r?\n){2,}', '', 内容)。 - 如果CSV中存在双引号包裹的合法多行字段,请调整替换逻辑,避免破坏字段结构——比如先判断行内是否有未闭合的引号,再决定是否合并换行。
内容的提问来源于stack exchange,提问作者bdoubleu
相关产品推荐
相关产品推荐

