如何在Python中基于两字符串列条件创建changeInLocation新列
Adding the
changeInLocation Column to Your Location Data Got it, let's walk through how to create the required changeInLocation column based on your specified rules.
Rule Recap
- Set to
0: When thepastLocationvalue exists anywhere in the comma-separatedcurrentLocationlist (including exact matches where both values are identical) - Set to
1: WhenpastLocationis not present incurrentLocationat all
Row-by-Row Validation
Let's verify each entry to make sure we're applying the rule correctly:
- Row 1:
delhivsbangalore→delhiisn't inbangalore→changeInLocation = 1 - Row 2:
delhivslondon,pune,delhi→delhiis explicitly in the list →changeInLocation = 0 - Row 3:
mumbaivsmumbai→ Exact match, somumbaiis present →changeInLocation = 0 - Row 4:
punevspune, noida→puneis the first entry in the list →changeInLocation = 0
Final Formatted Output
pastLocation | currentLocation | changeInLocation delhi | bangalore | 1 delhi | london,pune,delhi | 0 mumbai | mumbai | 0 pune | pune, noida | 0
Bonus: Python (Pandas) Implementation
If you're working with this data programmatically, here's a quick pandas snippet to generate the column automatically:
import pandas as pd # Initialize the data location_data = { "pastLocation": ["delhi", "delhi", "mumbai", "pune"], "currentLocation": ["bangalore", "london,pune,delhi", "mumbai", "pune, noida"] } df = pd.DataFrame(location_data) # Function to calculate the change value def get_location_change(past, current): # Split current locations, clean up whitespace, check for presence cleaned_current = [loc.strip() for loc in current.split(",")] return 0 if past in cleaned_current else 1 # Apply the function to create the new column df["changeInLocation"] = df.apply(lambda row: get_location_change(row["pastLocation"], row["currentLocation"]), axis=1) # Output in pipe-separated format print(df.to_csv(sep="|", index=False))
内容的提问来源于stack exchange,提问作者Mighty
相关产品推荐
相关产品推荐

