You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas按逗号拆分指定列字符串行的例外场景处理咨询

Handling Comma Splits in Pandas with Specific Exceptions

Great question! Since you only have three specific exceptions where commas shouldn't trigger splits, a practical and clean approach is to temporarily replace those "protected" commas with a unique placeholder, perform your split/explode workflow, then restore the commas afterward. Here's how to implement this:

Step-by-Step Solution

First, set up your sample data and define the exception phrases (the ones where commas should stay intact):

import pandas as pd

# Your sample DataFrame
df = pd.DataFrame({
    "col > 1": [
        "Personals, Financials, Data related to sexual preferences, sex life, and/or sexual orientation",
        "Personals, Financials",
        "Vendors, Procurement, subcontracting and vendor management"
    ]
})

# Define the exception phrases where commas should NOT be split
exception_phrases = [
    "Data related to sexual preferences, sex life, and/or sexual orientation",
    "Contract, salary and benefits",
    "Procurement, subcontracting and vendor management"
]

Next, process each target column (those with ">" in the name) using the placeholder trick:

# Loop through columns containing ">"
for col in df.columns[df.columns.str.contains(">")]:
    # Step 1: Replace commas in exception phrases with a unique placeholder
    temp_col = df[col].copy()
    for phrase in exception_phrases:
        # Swap commas in the exception for a rare placeholder
        temp_col = temp_col.str.replace(phrase, phrase.replace(",", "||COMMA||"), regex=False)
    
    # Step 2: Split on remaining commas and explode the values
    temp_col = temp_col.str.split(", ").explode().reset_index(drop=True)
    
    # Step 3: Restore the commas in exception phrases
    temp_col = temp_col.str.replace("||COMMA||", ",", regex=False)
    
    # Update the DataFrame with the processed column
    df = pd.DataFrame({col: temp_col})

Check the Result

Printing df will give you exactly your desired output:

col > 1
0                                                               Personals
1                                                              Financials
2  Data related to sexual preferences, sex life, and/or sexual orientation
3                                                               Personals
4                                                              Financials
5                                                                 Vendors
6                       Procurement, subcontracting and vendor management

Why This Approach Works

  • The unique placeholder (||COMMA||) avoids accidental matches with real data.
  • The logic stays focused and easy to maintain—just update the exception_phrases list if you need to add/remove exceptions later.
  • It preserves your original split/explode behavior for all commas outside the defined exceptions.

内容的提问来源于stack exchange,提问作者torkestativ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 14:52:26