You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将多个CSV文件按列合并?现有Python代码未达预期

问题描述

我需要把多个CSV文件按列合并成一个文件,但现有代码没达到预期效果。现有代码会把单个CSV的所有行合并成新CSV的一行,其他CSV作为新行添加。

我的代码如下:

import glob
import re
import csv
from bs4 import BeautifulSoup


for file in glob.glob('2023-02-10-upraveno/oh/pokus/??oh.csv'):
    raw_html = open(file)
    cleantext = BeautifulSoup(raw_html, "lxml").text 
    output = re.sub('\s+',' ', cleantext)      # saved the result using a variable
    print(output)                              # the variable can be reused
    row = [output]                             # as needed, in different contexts 
    with open('Ohs.csv', 'a') as csvFile:
        writer = csv.writer(csvFile)
        writer.writerow(row)

文件结构示例:

  • 文件1内容:
a,b
c,d
  • 文件2内容(同结构):
e,f
g,h

预期结果:

a,b,e,f,...
c,d,g,h,...

现有代码输出结果:

"a,b c,d"
"e,f g,h"

解决方案

问题核心是你把整个CSV文件的内容当成单行写入,且没有按行对应合并列。正确思路是先读取每个文件的所有行,再按行索引把对应行的内容拼接,最后写入新文件。

基础版代码(适用于纯CSV文件)

import glob
import csv

# 存储所有文件的行数据,每个元素对应一个文件的行列表
all_files_rows = []

# 遍历目标CSV文件
for file_path in glob.glob('2023-02-10-upraveno/oh/pokus/??oh.csv'):
    with open(file_path, 'r', newline='', encoding='utf-8') as f:
        reader = csv.reader(f)
        # 读取当前文件所有行并存储
        rows = list(reader)
        all_files_rows.append(rows)

# 按行索引合并列
merged_rows = []
if all_files_rows:
    # 以第一个文件的行数为基准(假设所有文件行数一致)
    num_rows = len(all_files_rows[0])
    for row_idx in range(num_rows):
        merged_row = []
        for file_rows in all_files_rows:
            # 拼接每个文件对应行的元素
            merged_row.extend(file_rows[row_idx])
        merged_rows.append(merged_row)

# 写入合并后的文件
with open('Ohs.csv', 'w', newline='', encoding='utf-8') as out_f:
    writer = csv.writer(out_f)
    writer.writerows(merged_rows)

带HTML清理的版本(适配原代码的HTML处理需求)

如果你的CSV文件包含HTML标签,先清理标签再解析CSV:

import glob
import csv
from bs4 import BeautifulSoup
import re

all_files_rows = []

for file_path in glob.glob('2023-02-10-upraveno/oh/pokus/??oh.csv'):
    with open(file_path, 'r', encoding='utf-8') as f:
        raw_html = f.read()
        # 移除HTML标签
        cleantext = BeautifulSoup(raw_html, "lxml").text
        # 把多余空白符替换为换行,恢复CSV行结构
        cleaned_content = re.sub(r'\s+', '\n', cleantext).strip()
        # 解析清理后的CSV内容
        reader = csv.reader(cleaned_content.splitlines())
        rows = list(reader)
        all_files_rows.append(rows)

# 合并列逻辑与基础版一致
if all_files_rows:
    num_rows = len(all_files_rows[0])
    merged_rows = []
    for row_idx in range(num_rows):
        merged_row = []
        for file_rows in all_files_rows:
            merged_row.extend(file_rows[row_idx])
        merged_rows.append(merged_row)

    with open('Ohs.csv', 'w', newline='', encoding='utf-8') as out_f:
        writer = csv.writer(out_f)
        writer.writerows(merged_rows)

内容的提问来源于stack exchange,提问作者Jan K

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 21:45:21