如何从DataFrame单元格的逗号分隔列表中提取连字符间字符并按规则转换格式?
Hey there! Let's solve this problem step by step. We need to process the Subset column in your DataFrame to extract the relevant parts from each comma-separated item, following the rules you outlined.
Sample Data First
Let's start with the sample DataFrame you provided to test our solutions:
import pandas as pd data = { "Name": ["Apple", "Bat"], "Subset": ["-AI-,-BI-A,-XC-,ZX-", "-po-,-IJ-,-IA-B"] } df = pd.DataFrame(data)
Method 1: Using a Helper Function (Readable)
This approach uses a simple function to process each entry, making it easy to follow and modify if needed.
Step-by-Step Function
We'll write a function that:
- Splits the string into individual items by commas
- Processes each item to remove the leading hyphen (if present) and everything after the next hyphen
- Joins the processed items back into a single string
def process_subset_entry(subset_str): # Split into individual items, stripping any extra whitespace items = [item.strip() for item in subset_str.split(',')] processed_items = [] for item in items: # Remove leading hyphen if it exists if item.startswith('-'): cleaned_item = item[1:] else: cleaned_item = item # Take everything up to the first hyphen (discard rest) final_part = cleaned_item.split('-')[0] processed_items.append(final_part) # Join back into a comma-separated string return ','.join(processed_items)
Apply the Function to the DataFrame
Now apply this function to the Subset column:
df['Processed_Subset'] = df['Subset'].apply(process_subset_entry)
Method 2: Using Regex (Concise & Fast)
If you prefer a more concise approach (great for large datasets), we can use pandas' string methods with regex to extract the desired parts in one line.
Regex Explanation
The regex pattern captures the relevant substring from each item:
(?<=^|,): Matches the start of the string or a comma (to identify the start of an item)\s*: Ignores any whitespace after commas-?: Accounts for an optional leading hyphen([^-]+): Captures one or more characters that aren't hyphens (this is the part we want)(?:-.*?)?(?=,|$): Ignores everything from the next hyphen to the end of the item
Code Implementation
df['Processed_Subset'] = df['Subset'].str.findall(r'(?<=^|,)\s*-?([^-]+)(?:-.*?)?(?=,|$)').str.join(',')
Result
Both methods will give you the same output:
| Name | Subset | Processed_Subset |
|---|---|---|
| Apple | -AI-,-BI-A,-XC-,ZX- | AI,BI,XC,ZX |
| Bat | -po-,-IJ-,-IA-B | po,IJ,IA |
Either method works perfectly—choose the one that fits your readability or performance needs!
内容的提问来源于stack exchange,提问作者spd

