You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从DataFrame单元格的逗号分隔列表中提取连字符间字符并按规则转换格式?

Solution to Extract Substrings from Comma-Separated List in DataFrame

Hey there! Let's solve this problem step by step. We need to process the Subset column in your DataFrame to extract the relevant parts from each comma-separated item, following the rules you outlined.

Sample Data First

Let's start with the sample DataFrame you provided to test our solutions:

import pandas as pd

data = {
    "Name": ["Apple", "Bat"],
    "Subset": ["-AI-,-BI-A,-XC-,ZX-", "-po-,-IJ-,-IA-B"]
}
df = pd.DataFrame(data)

Method 1: Using a Helper Function (Readable)

This approach uses a simple function to process each entry, making it easy to follow and modify if needed.

Step-by-Step Function

We'll write a function that:

  1. Splits the string into individual items by commas
  2. Processes each item to remove the leading hyphen (if present) and everything after the next hyphen
  3. Joins the processed items back into a single string
def process_subset_entry(subset_str):
    # Split into individual items, stripping any extra whitespace
    items = [item.strip() for item in subset_str.split(',')]
    processed_items = []
    
    for item in items:
        # Remove leading hyphen if it exists
        if item.startswith('-'):
            cleaned_item = item[1:]
        else:
            cleaned_item = item
        
        # Take everything up to the first hyphen (discard rest)
        final_part = cleaned_item.split('-')[0]
        processed_items.append(final_part)
    
    # Join back into a comma-separated string
    return ','.join(processed_items)

Apply the Function to the DataFrame

Now apply this function to the Subset column:

df['Processed_Subset'] = df['Subset'].apply(process_subset_entry)

Method 2: Using Regex (Concise & Fast)

If you prefer a more concise approach (great for large datasets), we can use pandas' string methods with regex to extract the desired parts in one line.

Regex Explanation

The regex pattern captures the relevant substring from each item:

  • (?<=^|,): Matches the start of the string or a comma (to identify the start of an item)
  • \s*: Ignores any whitespace after commas
  • -?: Accounts for an optional leading hyphen
  • ([^-]+): Captures one or more characters that aren't hyphens (this is the part we want)
  • (?:-.*?)?(?=,|$): Ignores everything from the next hyphen to the end of the item

Code Implementation

df['Processed_Subset'] = df['Subset'].str.findall(r'(?<=^|,)\s*-?([^-]+)(?:-.*?)?(?=,|$)').str.join(',')

Result

Both methods will give you the same output:

NameSubsetProcessed_Subset
Apple-AI-,-BI-A,-XC-,ZX-AI,BI,XC,ZX
Bat-po-,-IJ-,-IA-Bpo,IJ,IA

Either method works perfectly—choose the one that fits your readability or performance needs!

内容的提问来源于stack exchange,提问作者spd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 15:27:27