如何使用Pandas将文本文件的列转换为单行CSV文件?
Using Pandas to Convert Variable-Length Text Lines to Column-Wise Single-Line CSV
I see you're trying to convert a text file with variable-length rows into two single-line CSV entries, ordered column-wise. Your initial code ran into issues because zip() stops at the shortest line, and you referenced undefined columns. Let's fix this with Pandas, which handles variable-length data nicely.
Step-by-Step Solution:
- Read and Split Input Lines: First, we read the text file and split each line into individual elements.
- Split into Groups: Separate the data string lines and number lines into two distinct groups.
- Process Each Group with Pandas:
- Convert each group into a DataFrame (Pandas will automatically fill missing values with
NaNfor shorter rows). - Iterate over each column, collect non-null values, and concatenate them in column-wise order.
- Join the collected values into a comma-separated string.
- Convert each group into a DataFrame (Pandas will automatically fill missing values with
- Write to CSV: Save the two processed lines into your output CSV file.
Full Code:
import pandas as pd # Read the input text file with open("file.txt", "r") as f: # Split each line into elements, skipping any empty lines lines = [line.strip().split() for line in f if line.strip()] # Split into data group (first 4 lines) and number group (next 4 lines) data_group = lines[:4] num_group = lines[4:] # Process the data string group df_data = pd.DataFrame(data_group) data_result = [] for col in df_data.columns: # Collect non-null values from each column data_result.extend(df_data[col].dropna().tolist()) data_line = ','.join(data_result) # Process the number group df_num = pd.DataFrame(num_group) num_result = [] for col in df_num.columns: # Convert numbers to strings and collect non-null values num_result.extend(df_num[col].dropna().astype(str).tolist()) num_line = ','.join(num_result) # Write the result to CSV with open("file.csv", "w") as fout: fout.write(f"{data_line}\n") fout.write(f"{num_line}\n")
How This Works:
- DataFrame Handling: When we create a DataFrame from variable-length rows, Pandas fills missing positions with
NaN, which makes it easy to filter out empty entries later. - Column-Wise Collection: By iterating over each column and collecting non-null values, we ensure we follow the exact order you want: all first-column elements, then all second-column elements, and so on.
- String Conversion: For the number group, we convert values to strings to ensure they join correctly into a CSV line.
This will produce the exact output you described, with each group's elements ordered column-wise in a single CSV line.
内容的提问来源于stack exchange,提问作者Sandeep Singh
相关产品推荐
相关产品推荐

