You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在不使用Pandas的情况下用Python合并两个CSV文件?

Solution: Merge Two CSVs Without Pandas

Hey there! Let's work through this step by step. Your current code has a couple of key issues (wrong delimiter, exhausted reader from nested loops) but we can fix this by using a lookup dictionary to efficiently match rows by ID.

First, Let's Break Down the Problem

We need to:

  1. Match rows from the second CSV (xyz_id) to rows in the first CSV (id)
  2. Combine each matching pair into the output format you specified
  3. Avoid inefficient nested loops that break the CSV reader

Step-by-Step Fix & Full Code

Here's a complete working solution that addresses your requirements:

import csv
import sys

# File paths (update these if your paths are different)
first_csv = '/home/user/Downloads/FW__Json_FIles/withoutpanda.csv'
second_csv = '/home/user/Downloads/FW__Json_FIles/forms.csv'

# Step 1: Build a lookup dictionary from the first CSV (using id as key)
id_lookup = {}
with open(first_csv, 'r') as f:
    reader = csv.reader(f, delimiter=',')  # Fix: Use comma delimiter, not |
    first_header = next(reader)  # Capture the header row
    for row in reader:
        row_id = row[3]  # id is the 4th column (index 3)
        id_lookup[row_id] = row  # Store the full row for quick lookup

# Step 2: Process the second CSV and merge with matching rows from first CSV
with open(second_csv, 'r') as f:
    reader = csv.reader(f, delimiter=',')  # Fix: Use comma delimiter here too
    second_header = next(reader)  # Capture the second CSV's header
    # Combine headers to match your desired output
    combined_header = second_header + first_header
    print(','.join(combined_header))

    for row in reader:
        xyz_id = row[0]
        # Look up the matching row from the first CSV
        if xyz_id in id_lookup:
            matching_row = id_lookup[xyz_id]
            # Combine the two rows and print
            combined_row = row + matching_row
            print(','.join(combined_row))
        else:
            # Optional: Handle cases where no matching ID exists
            print(f"Warning: No matching ID found for xyz_id {xyz_id}", file=sys.stderr)

Key Fixes & Explanations

  1. Delimiter Correction: Your sample CSVs use commas, but your original code used | as the delimiter. This was causing the reader to parse each line as a single column instead of splitting values correctly.
  2. Lookup Dictionary: By storing rows from the first CSV in a dictionary with id as the key, we avoid nested loops (which would exhaust the second CSV reader after the first iteration) and enable O(1) fast lookups.
  3. Header Handling: We explicitly capture headers from both files and combine them to match your desired output structure.
  4. Error Resilience: Added a warning for cases where an xyz_id in the second CSV has no matching id in the first (you can remove this if you don't need it).

What Was Wrong With Your Original Code?

  • Nested Loops: The reader2 gets fully consumed on the first iteration of row1, so there's no data left for subsequent rows in the first CSV.
  • Incorrect Delimiter: Using | instead of , meant the CSV reader wasn't splitting your data into columns properly.

内容的提问来源于stack exchange,提问作者Rakshit Sakhuja

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:24:03