Python实现长格式转宽格式及数据分批拆分、列表转换咨询
Hey there! Let's break down your questions one by one with practical, runnable code examples that you can test right away:
The go-to tool for this is the pandas library—it's designed exactly for this kind of data reshaping.
First, quick clarification: long format data has one row per observation (e.g., each row tracks a single attribute for an entity), while wide format has one row per entity with attributes as columns.
Example with pivot() (for unique index-column pairs):
import pandas as pd # Sample long-format data long_df = pd.DataFrame({ 'user_id': ['user1', 'user1', 'user2', 'user2'], 'metric': ['age', 'income', 'age', 'income'], 'value': [28, 75000, 35, 90000] }) # Convert to wide format wide_df = long_df.pivot( index='user_id', # Column to use as rows in wide format columns='metric', # Column to spread into new columns values='value' # Values to fill the new columns ).reset_index() # Clean up the column name (optional but nicer for readability) wide_df.columns.name = None print(wide_df)
Output:
user_id age income 0 user1 28 75000 1 user2 35 90000
If you have duplicate index-column pairs:
Use pivot_table() with an aggregation function (like mean, sum, or first) to handle duplicates gracefully:
# Long data with duplicate entries long_df_duplicates = pd.DataFrame({ 'user_id': ['user1', 'user1', 'user1', 'user2'], 'metric': ['income', 'income', 'age', 'income'], 'value': [75000, 78000, 28, 90000] }) wide_df = long_df_duplicates.pivot_table( index='user_id', columns='metric', values='value', aggfunc='mean' # Average duplicate income values ).reset_index() wide_df.columns.name = None print(wide_df)
Let's tackle this step by step:
Step 1: Convert your raw string to a list
First, split your space-separated string into individual items:
# Your raw data string raw_data = "A292340 A291630 A278240 A267770 A267490 A261250 A261110 A253150 A252400 A253250 A243890 A243880 A236350 A233740 A233160 A225800 A225060 A225050 A225040 A225130 A219900 A204450 A204480 A204420 A196030 A196220 A167860 A152500 A123320 A122630" # Split into a list of individual items data_list = raw_data.split() print(f"Total items: {len(data_list)}") # This is exactly 30, so one group here
Step 2: Split into groups of 30
Use a list comprehension to chunk the list into groups of 30. This works seamlessly even if your data is longer than 30 items:
group_size = 30 data_groups = [data_list[i:i+group_size] for i in range(0, len(data_list), group_size)] # Verify the groups for idx, group in enumerate(data_groups, 1): print(f"Group {idx} has {len(group)} items: {group[:5]}...") # Print first 5 items of each group
Step 3: Save each group to a file
You can save each group to a separate text file (adjust to CSV if needed):
for idx, group in enumerate(data_groups, 1): with open(f"data_group_{idx}.txt", 'w') as file: file.write('\n'.join(group)) # Write each item on a new line print(f"Saved Group {idx} to data_group_{idx}.txt")
Converting print output to a list
If you're asking how to collect content you'd normally print into a list, the simplest approach is to build the list directly instead of printing first:
# Example: Collect items into a list while still printing collected_list = [] for item in data_list: collected_list.append(item) # Append to the list print(item) # Optional: keep printing if you need to see the output # Now collected_list contains all the items you printed print("\nCollected list:", collected_list)
If you ever need to capture actual print output from a function that only prints (instead of returning data), you can use io.StringIO, but building the list directly is cleaner and more efficient for most use cases.
内容的提问来源于stack exchange,提问作者Hannah Lee

