Pandas:如何依据source与website的匹配修改DataFrame的type列值
Hey there! Let's walk through how to modify the 'type' column based on whether the 'source' value exists within the 'website' string. Here's a straightforward, efficient approach using pandas and numpy:
First, let's set up our original DataFrame as you provided:
import pandas as pd import numpy as np raw_data = { 'website': ['bbc.com', 'cnn.com', 'google.com', 'facebook.com'], 'type': ['image', 'audio', 'image', 'video'], 'source': ['bbc','google','stackoverflow','facebook'] } df = pd.DataFrame(raw_data, columns=['website', 'type', 'source'])
Next, we’ll create a row-wise condition to check if each 'source' is a substring of its corresponding 'website'. Then we’ll use this condition to append the correct suffix to the 'type' column:
# Check if source exists in website for each row is_first_party = df.apply(lambda row: row['source'] in row['website'], axis=1) # Update the type column with the appropriate suffix df['type'] = np.where(is_first_party, df['type'] + '_1stParty', df['type'] + '_3rdParty')
If you print the updated DataFrame now, you’ll get exactly the result you’re looking for:
website type source 0 bbc.com image_1stParty bbc 1 cnn.com audio_3rdParty google 2 google.com image_3rdParty stackoverflow 3 facebook.com video_1stParty facebook
This method uses pandas' vectorized operations where possible, keeping things efficient even for larger datasets, and the lambda function handles the row-specific check cleanly.
内容的提问来源于stack exchange,提问作者Saad

