如何用Python合并指定条数的CSV数据?解决代码合并全量数据问题
How to Merge Specific Ranges of Rows from Two CSV Files with Pandas
Hey there! Let's break down why your current code isn't giving you the result you want, and fix it to extract exactly the rows you need (first 10, then 11-20 from trainNegatif.csv) and merge them sequentially into testNegatif.csv.
First, Let's Diagnose the Issue
Your current code uses nrows=10 to read the first 10 rows of trainNegatif.csv, but you mentioned it's merging all data instead. A few possible reasons for this:
- Maybe you accidentally ran an older version of the code that read the full file (check your code history or restart your kernel if using Jupyter).
- Or, perhaps you intended to merge both the first 10 and 11-20 rows, but your current code only handles the first 10—so you're missing the second batch, making it feel like nothing changed.
The Fix: Extract Specific Row Ranges & Merge Sequentially
Here's a reliable way to get the exact rows you need and merge them properly. We'll use pd.concat() instead of append() (since append() is deprecated in newer pandas versions):
import pandas as pd # Read the base CSV (testNegatif.csv) df_base = pd.read_csv('testNegatif.csv') # Extract first 10 rows from trainNegatif.csv df_first_10 = pd.read_csv('trainNegatif.csv', nrows=10) # Extract rows 11-20 from trainNegatif.csv (two methods below) # Method 1: Read the full file then slice (best for small to medium files) df_full_train = pd.read_csv('trainNegatif.csv') df_11_to_20 = df_full_train.iloc[10:20] # pandas uses 0-indexing, so 10-19 = rows 11-20 # Method 2: Skip first 10 data rows during read (good for large files to save memory) # Note: range(1,11) skips the first 10 data rows (row 0 is the header) # df_11_to_20 = pd.read_csv('trainNegatif.csv', skiprows=range(1, 11), nrows=10) # Merge sequentially: base → first 10 → 11-20 merged_df = pd.concat([df_base, df_first_10, df_11_to_20], ignore_index=True) # Save the final output merged_df.to_csv("output.csv", sep=',', index=False)
Key Notes:
iloc[10:20]selects rows from index 10 to 19 (inclusive), which corresponds to the 11th to 20th rows in your CSV (since pandas uses 0-based indexing).ignore_index=Trueresets the index of the merged dataframe so you don't have duplicate index values.- Adding
index=Falsetoto_csv()prevents pandas from writing an extra index column to your output file.
内容的提问来源于stack exchange,提问作者Trya Sovi Kartikasari
相关产品推荐
相关产品推荐

