You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现长格式转宽格式及数据分批拆分、列表转换咨询

Hey there! Let's break down your questions one by one with practical, runnable code examples that you can test right away:

1. Converting Long Format Data to Wide Format in Python

The go-to tool for this is the pandas library—it's designed exactly for this kind of data reshaping.

First, quick clarification: long format data has one row per observation (e.g., each row tracks a single attribute for an entity), while wide format has one row per entity with attributes as columns.

Example with pivot() (for unique index-column pairs):

import pandas as pd

# Sample long-format data
long_df = pd.DataFrame({
    'user_id': ['user1', 'user1', 'user2', 'user2'],
    'metric': ['age', 'income', 'age', 'income'],
    'value': [28, 75000, 35, 90000]
})

# Convert to wide format
wide_df = long_df.pivot(
    index='user_id',  # Column to use as rows in wide format
    columns='metric', # Column to spread into new columns
    values='value'    # Values to fill the new columns
).reset_index()

# Clean up the column name (optional but nicer for readability)
wide_df.columns.name = None

print(wide_df)

Output:

user_id  age  income
0   user1   28   75000
1   user2   35   90000

If you have duplicate index-column pairs:

Use pivot_table() with an aggregation function (like mean, sum, or first) to handle duplicates gracefully:

# Long data with duplicate entries
long_df_duplicates = pd.DataFrame({
    'user_id': ['user1', 'user1', 'user1', 'user2'],
    'metric': ['income', 'income', 'age', 'income'],
    'value': [75000, 78000, 28, 90000]
})

wide_df = long_df_duplicates.pivot_table(
    index='user_id',
    columns='metric',
    values='value',
    aggfunc='mean'  # Average duplicate income values
).reset_index()

wide_df.columns.name = None
print(wide_df)
2. Splitting Long Data into Groups of 30 & Saving, Plus Converting Print Output to a List

Let's tackle this step by step:

Step 1: Convert your raw string to a list

First, split your space-separated string into individual items:

# Your raw data string
raw_data = "A292340 A291630 A278240 A267770 A267490 A261250 A261110 A253150 A252400 A253250 A243890 A243880 A236350 A233740 A233160 A225800 A225060 A225050 A225040 A225130 A219900 A204450 A204480 A204420 A196030 A196220 A167860 A152500 A123320 A122630"

# Split into a list of individual items
data_list = raw_data.split()
print(f"Total items: {len(data_list)}")  # This is exactly 30, so one group here

Step 2: Split into groups of 30

Use a list comprehension to chunk the list into groups of 30. This works seamlessly even if your data is longer than 30 items:

group_size = 30
data_groups = [data_list[i:i+group_size] for i in range(0, len(data_list), group_size)]

# Verify the groups
for idx, group in enumerate(data_groups, 1):
    print(f"Group {idx} has {len(group)} items: {group[:5]}...")  # Print first 5 items of each group

Step 3: Save each group to a file

You can save each group to a separate text file (adjust to CSV if needed):

for idx, group in enumerate(data_groups, 1):
    with open(f"data_group_{idx}.txt", 'w') as file:
        file.write('\n'.join(group))  # Write each item on a new line
    print(f"Saved Group {idx} to data_group_{idx}.txt")

Converting print output to a list

If you're asking how to collect content you'd normally print into a list, the simplest approach is to build the list directly instead of printing first:

# Example: Collect items into a list while still printing
collected_list = []
for item in data_list:
    collected_list.append(item)  # Append to the list
    print(item)  # Optional: keep printing if you need to see the output

# Now collected_list contains all the items you printed
print("\nCollected list:", collected_list)

If you ever need to capture actual print output from a function that only prints (instead of returning data), you can use io.StringIO, but building the list directly is cleaner and more efficient for most use cases.


内容的提问来源于stack exchange,提问作者Hannah Lee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:42:17