如何将多个CSV文件按列合并?现有Python代码未达预期
问题描述
我需要把多个CSV文件按列合并成一个文件,但现有代码没达到预期效果。现有代码会把单个CSV的所有行合并成新CSV的一行,其他CSV作为新行添加。
我的代码如下:
import glob import re import csv from bs4 import BeautifulSoup for file in glob.glob('2023-02-10-upraveno/oh/pokus/??oh.csv'): raw_html = open(file) cleantext = BeautifulSoup(raw_html, "lxml").text output = re.sub('\s+',' ', cleantext) # saved the result using a variable print(output) # the variable can be reused row = [output] # as needed, in different contexts with open('Ohs.csv', 'a') as csvFile: writer = csv.writer(csvFile) writer.writerow(row)
文件结构示例:
- 文件1内容:
a,b c,d
- 文件2内容(同结构):
e,f g,h
预期结果:
a,b,e,f,... c,d,g,h,...
现有代码输出结果:
"a,b c,d" "e,f g,h"
解决方案
问题核心是你把整个CSV文件的内容当成单行写入,且没有按行对应合并列。正确思路是先读取每个文件的所有行,再按行索引把对应行的内容拼接,最后写入新文件。
基础版代码(适用于纯CSV文件)
import glob import csv # 存储所有文件的行数据,每个元素对应一个文件的行列表 all_files_rows = [] # 遍历目标CSV文件 for file_path in glob.glob('2023-02-10-upraveno/oh/pokus/??oh.csv'): with open(file_path, 'r', newline='', encoding='utf-8') as f: reader = csv.reader(f) # 读取当前文件所有行并存储 rows = list(reader) all_files_rows.append(rows) # 按行索引合并列 merged_rows = [] if all_files_rows: # 以第一个文件的行数为基准(假设所有文件行数一致) num_rows = len(all_files_rows[0]) for row_idx in range(num_rows): merged_row = [] for file_rows in all_files_rows: # 拼接每个文件对应行的元素 merged_row.extend(file_rows[row_idx]) merged_rows.append(merged_row) # 写入合并后的文件 with open('Ohs.csv', 'w', newline='', encoding='utf-8') as out_f: writer = csv.writer(out_f) writer.writerows(merged_rows)
带HTML清理的版本(适配原代码的HTML处理需求)
如果你的CSV文件包含HTML标签,先清理标签再解析CSV:
import glob import csv from bs4 import BeautifulSoup import re all_files_rows = [] for file_path in glob.glob('2023-02-10-upraveno/oh/pokus/??oh.csv'): with open(file_path, 'r', encoding='utf-8') as f: raw_html = f.read() # 移除HTML标签 cleantext = BeautifulSoup(raw_html, "lxml").text # 把多余空白符替换为换行,恢复CSV行结构 cleaned_content = re.sub(r'\s+', '\n', cleantext).strip() # 解析清理后的CSV内容 reader = csv.reader(cleaned_content.splitlines()) rows = list(reader) all_files_rows.append(rows) # 合并列逻辑与基础版一致 if all_files_rows: num_rows = len(all_files_rows[0]) merged_rows = [] for row_idx in range(num_rows): merged_row = [] for file_rows in all_files_rows: merged_row.extend(file_rows[row_idx]) merged_rows.append(merged_row) with open('Ohs.csv', 'w', newline='', encoding='utf-8') as out_f: writer = csv.writer(out_f) writer.writerows(merged_rows)
内容的提问来源于stack exchange,提问作者Jan K
相关产品推荐
相关产品推荐

