You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas将按行存储的字符串数据转换为7列CSV?

Convert Your Line-by-Line Formatted File to CSV with Pandas

Hey there! I see you’ve already nailed a Bash solution using loops and modulo, but let’s switch over to Pandas for a cleaner, more Pythonic approach. Here’s how to transform that odd line-per-field file into a proper 7-column CSV:

Step-by-Step Solution

Let’s adjust your existing code to process the data correctly with Pandas:

import sys
import argparse
import pandas as pd

parser = argparse.ArgumentParser()
parser.add_argument('file', help="this is the file you want to open")
args = parser.parse_args()
print("file name:", args.file)

# Read and clean the file content
with open(args.file, 'r') as f:
    # Strip newline characters and filter out empty lines to avoid data gaps
    lines = [line.strip() for line in f.readlines() if line.strip()]

# Extract headers (first 7 lines of the cleaned content)
headers = lines[:7]

# Group remaining lines into chunks of 7 to form data rows
data = [lines[i:i+7] for i in range(7, len(lines), 7)]

# Convert to DataFrame and save as CSV
df = pd.DataFrame(data, columns=headers)
df.to_csv('output.csv', index=False)

print(f"Success! Your CSV is saved as 'output.csv'")

How This Works

  • Cleaning the Input: We use strip() to remove newline characters and extra whitespace, then filter out empty lines to prevent missing values or misaligned rows.
  • Extracting Headers: The first 7 lines are our column names, so we grab lines[:7] directly.
  • Grouping Data: Using a list comprehension with range(7, len(lines), 7), we slice the remaining content into chunks of 7 lines—each chunk becomes a single row in our DataFrame.
  • Saving to CSV: Pandas handles all the CSV formatting heavy lifting, and index=False ensures we don’t add an unnecessary index column to the output.

Quick Example

If your input looks like this (truncated):

Filename
Dataset
Level
Duration
Accuracy
Speed Ratio
Completed
file_001.mp3
DatasetA
Level1
120
98.5
1.0
Yes
file_002.mp3
DatasetB
Level2
90
95.2
0.8
No

The resulting CSV will have clean columns (Filename, Dataset, Level, Duration, Accuracy, Speed Ratio, Completed) with each row holding the corresponding values for each file.

This approach is concise, easy to tweak later, and plays to Pandas’ strengths for tabular data handling.

内容的提问来源于stack exchange,提问作者Morris Gevirtz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:00:16