Python中按多自定义分隔符拆分列:解析Creative列的结构化数据需求
Hey there! Let's fix that column splitting issue you're having. First, let's break down why your original code threw an error, then jump to a clean, scalable solution.
Why your original code failed
The error missing ), unterminated subpattern at position 2 happens because when you use str.split('cs('), pandas treats the input as a regular expression by default. In regex, ( is a special character that starts a subpattern—since you didn't escape it or close it, the parser gets confused. Even if you fixed that, splitting on just cs( only handles one part of your data, which isn't ideal for all the key-value pairs in your Creative column.
The better approach: Extract all key-value pairs with regex
Your Creative column follows a consistent pattern: XX(Value) where XX is a 2-letter column name. We can use regex to extract all these key-value pairs automatically, then convert them into a new DataFrame. Here's how:
Step-by-step code
import pandas as pd import re # Example DataFrame (replace with your actual data) df = pd.DataFrame({ 'Creative': [ 'pn(2021)io(302)ta(Yes)pt(Blue)cn(John)cs(Doe)', 'pn(2023)io(501)ta(Yes)pt(Green)cn(Bob)cs(White)' ] }) # Regex pattern to match 2-letter keys and their values in parentheses # - ([a-z]{2}): Captures the 2-letter key (lowercase, adjust if needed for uppercase) # - \(.*?\): Captures the value inside parentheses (non-greedy to stop at the first ')') pattern = re.compile(r'([a-z]{2})\((.*?)\)') # For each row, extract all key-value pairs into a dict, then convert to a DataFrame split_columns = df['Creative'].apply(lambda x: dict(pattern.findall(x))).apply(pd.Series) # Merge the split columns back with your original DataFrame (optional, if you want to keep the original Creative column) final_df = pd.concat([df, split_columns], axis=1) # If you don't need the original Creative column anymore, just use split_columns directly # final_df = split_columns print(final_df)
Output
This will give you exactly the format you want:
Creative pn io ta pt cn cs 0 pn(2021)io(302)ta(Yes)pt(Blue)cn(John)cs(Doe) 2021 302 Yes Blue John Doe 1 pn(2023)io(501)ta(Yes)pt(Green)cn(Bob)cs(White) 2023 501 Yes Green Bob White
Why this works
- Automatically handles all keys: You don't have to hardcode each key like
cs(orpn(—the regex finds every 2-letter code followed by parentheses. - Scalable: If your Creative column ever adds new 2-letter keys, this code will automatically include them as new columns without any changes.
- Clean and readable: It turns the messy string into structured data in one go.
Quick fix for your original code (if you really wanted to split on cs()
If you just wanted to split on cs( specifically, you can tell pandas to treat the string as literal text instead of regex by adding regex=False:
df[['Creative', 'cs']] = df['Creative'].str.split('cs(', expand=True, regex=False) df['cs'] = df['cs'].str.replace(')', '')
But remember, this only splits out the cs value—you'd have to repeat this for every other key, which is tedious. The regex method is way better for your use case.
内容的提问来源于stack exchange,提问作者partridge20

