如何简洁拆分pandas DataFrame单列为多列?
Hey there! Totally get wanting a cleaner solution than a tedious workaround—pandas has some great built-in tools for this. Let's cover the two most common scenarios, depending on your column's format:
Case 1: Fixed Separator (e.g., comma, space, hyphen)
If your column values are separated by a consistent character (like a,b,c or x y z), the str.split() method with expand=True is your go-to. It directly splits the column into multiple new columns in one line.
Example Code:
Suppose your DataFrame looks like this:
import pandas as pd df = pd.DataFrame({'combined_col': ['apple,banana,cherry', 'dog,cat,bird', 'red,green,blue']})
Split into 3 columns in a single step:
df[['col1', 'col2', 'col3']] = df['combined_col'].str.split(',', expand=True)
Just replace ',' with your actual separator (like ' ' for spaces) and adjust the column names to match your needs.
Case 2: Pattern-Based Split (Regex)
If your column has a more structured but variable format (like item1:value1 | item2:value2 | item3:value3 or mixed delimiters), use str.extract() with a regex pattern to capture each segment precisely.
Example Code:
For a column formatted like '2023-10-05 | New York | $50', extract each part with:
df[['date', 'city', 'price']] = df['combined_col'].str.extract(r'(.*) \| (.*) \| (.*)')
The regex (.*) captures each group separated by |—tweak the pattern to match your specific data structure.
Bonus: Handling Uneven Splits
If some rows have fewer than 3 elements, add n=3 to limit splits to 3 parts, and pandas will automatically fill missing values with NaN:
df[['col1', 'col2', 'col3']] = df['combined_col'].str.split(',', n=3, expand=True)
Both methods eliminate the need for manual row-by-row processing—way more concise and pandas-idiomatic! Hope one of these fits your use case.
内容的提问来源于stack exchange,提问作者arnaud

