如何用Pandas将按行存储的字符串数据转换为7列CSV?
Convert Your Line-by-Line Formatted File to CSV with Pandas
Hey there! I see you’ve already nailed a Bash solution using loops and modulo, but let’s switch over to Pandas for a cleaner, more Pythonic approach. Here’s how to transform that odd line-per-field file into a proper 7-column CSV:
Step-by-Step Solution
Let’s adjust your existing code to process the data correctly with Pandas:
import sys import argparse import pandas as pd parser = argparse.ArgumentParser() parser.add_argument('file', help="this is the file you want to open") args = parser.parse_args() print("file name:", args.file) # Read and clean the file content with open(args.file, 'r') as f: # Strip newline characters and filter out empty lines to avoid data gaps lines = [line.strip() for line in f.readlines() if line.strip()] # Extract headers (first 7 lines of the cleaned content) headers = lines[:7] # Group remaining lines into chunks of 7 to form data rows data = [lines[i:i+7] for i in range(7, len(lines), 7)] # Convert to DataFrame and save as CSV df = pd.DataFrame(data, columns=headers) df.to_csv('output.csv', index=False) print(f"Success! Your CSV is saved as 'output.csv'")
How This Works
- Cleaning the Input: We use
strip()to remove newline characters and extra whitespace, then filter out empty lines to prevent missing values or misaligned rows. - Extracting Headers: The first 7 lines are our column names, so we grab
lines[:7]directly. - Grouping Data: Using a list comprehension with
range(7, len(lines), 7), we slice the remaining content into chunks of 7 lines—each chunk becomes a single row in our DataFrame. - Saving to CSV: Pandas handles all the CSV formatting heavy lifting, and
index=Falseensures we don’t add an unnecessary index column to the output.
Quick Example
If your input looks like this (truncated):
Filename Dataset Level Duration Accuracy Speed Ratio Completed file_001.mp3 DatasetA Level1 120 98.5 1.0 Yes file_002.mp3 DatasetB Level2 90 95.2 0.8 No
The resulting CSV will have clean columns (Filename, Dataset, Level, Duration, Accuracy, Speed Ratio, Completed) with each row holding the corresponding values for each file.
This approach is concise, easy to tweak later, and plays to Pandas’ strengths for tabular data handling.
内容的提问来源于stack exchange,提问作者Morris Gevirtz
相关产品推荐
相关产品推荐

