如何用Python/Pandas实现列转行及处理机器学习数据集标签映射
Hey there! Let's break down your two data processing questions with Pandas and get you sorted out.
If you're looking to split text within a single column into separate rows (like when each cell has multiple values separated by a delimiter such as commas or spaces), the explode() method paired with str.split() is your go-to tool. Here's a concrete example:
Example Code
import pandas as pd # Sample DataFrame with a single text column df = pd.DataFrame({ 'text': ['apple, banana, cherry', 'dog; cat; rabbit', 'sun moon stars'] }) # Split the text by delimiter (adjust to match your data's separator) and explode into rows df['text'] = df['text'].str.split(', ') # Swap with '; ' or ' ' depending on your data result_df = df.explode('text').reset_index(drop=True) print(result_df)
Output
text 0 apple 1 banana 2 cherry 3 dog 4 cat 5 rabbit 6 sun 7 moon 8 stars
If your column already has individual text entries that just need to be converted into a single column of rows (no splitting required), explode() still works—just wrap each value in a list first:
df['text'] = df['text'].apply(lambda x: [x]) result_df = df.explode('text')
Got it, let's tackle your machine learning dataset task. From your description:
sol = 0meanssent0violates common sense (label 0), sosent1is valid (label 1)- I’m assuming
sol = 1meanssent1violates common sense (label 0), sosent0is valid (label 1)
Here are two straightforward methods to achieve this:
Method 1: Use melt() to Reshape and Add Labels
This method converts your wide-form DataFrame to long-form first, then applies conditional logic to assign labels.
import pandas as pd # Sample dataset matching your structure df = pd.DataFrame({ 'sent0': ['The sun rises in the west', 'Water boils at 100°C', 'Plants need sunlight'], 'sent1': ['The sun rises in the east', 'Water boils at 0°C', 'Plants need darkness'], 'sol': [0, 1, 1] }) # Reshape the DataFrame to long format melted = df.melt( id_vars='sol', value_vars=['sent0', 'sent1'], var_name='sentence_source', value_name='sentence' ) # Assign labels based on sol and sentence source melted['label'] = melted.apply( lambda row: 0 if (row['sol'] == 0 and row['sentence_source'] == 'sent0') or (row['sol'] == 1 and row['sentence_source'] == 'sent1') else 1, axis=1 ) # Keep only the columns you need final_df = melted[['sentence', 'label']].reset_index(drop=True) print(final_df)
Method 2: Process Columns Separately and Concatenate
If you prefer a more explicit approach, handle sent0 and sent1 individually then combine them.
import pandas as pd # Same sample dataset df = pd.DataFrame({ 'sent0': ['The sun rises in the west', 'Water boils at 100°C', 'Plants need sunlight'], 'sent1': ['The sun rises in the east', 'Water boils at 0°C', 'Plants need darkness'], 'sol': [0, 1, 1] }) # Process sent0: label 0 if sol=0, else 1 sent0_df = df.rename(columns={'sent0': 'sentence'}) sent0_df['label'] = sent0_df['sol'].apply(lambda x: 0 if x == 0 else 1) # Process sent1: label 1 if sol=0, else 0 sent1_df = df.rename(columns={'sent1': 'sentence'}) sent1_df['label'] = sent1_df['sol'].apply(lambda x: 1 if x == 0 else 0) # Combine both DataFrames final_df = pd.concat([sent0_df, sent1_df], ignore_index=True)[['sentence', 'label']] print(final_df)
Output for Both Methods
sentence label 0 The sun rises in the west 0 1 Water boils at 100°C 1 2 Plants need sunlight 1 3 The sun rises in the east 1 4 Water boils at 0°C 0 5 Plants need darkness 0
Either method will get you the merged column with correct labels—pick whichever makes more sense for your workflow!
内容的提问来源于stack exchange,提问作者LeoCoder444

