如何基于Key列动态提取Pandas DataFrame中Vector列的列表元素
Here are two practical ways to create the Value column using the Key column as the dynamic index for the Vector list:
Method 1: Use df.apply() with axis=1
This approach uses apply() with axis=1 to access each row's values directly. The lambda function retrieves the element from Vector at the index specified by Key:
import pandas as pd x1 = ['A1','B1','C1'] x2 = ['A2','B2','C2','D2'] df = pd.DataFrame([['ID_1',0,x1],['ID_2',2,x2]], columns=['ID','Key','Vector']) # Add the Value column df['Value'] = df.apply(lambda row: row['Vector'][row['Key']], axis=1) print(df)
Output:
ID Key Vector Value 0 ID_1 0 [A1, B1, C1] A1 1 ID_2 2 [A2, B2, C2, D2] C2
Method 2: List Comprehension (Faster for Large Data)
List comprehensions are more efficient than apply() for large datasets. Pair each Vector with its corresponding Key using zip():
df['Value'] = [vec[key] for vec, key in zip(df['Vector'], df['Key'])]
This produces the same result but with better performance when working with larger DataFrames.
Edge Case Handling
If Key values might exceed the length of the corresponding Vector list (causing an IndexError), add error handling to return a default value like None:
# Apply with error handling df['Value'] = df.apply(lambda row: row['Vector'][row['Key']] if row['Key'] < len(row['Vector']) else None, axis=1) # List comprehension with error handling df['Value'] = [vec[key] if key < len(vec) else None for vec, key in zip(df['Vector'], df['Key'])]
内容的提问来源于stack exchange,提问作者guedj

