如何利用另一张表中元组元素的频率生成复合频率新列?
new_col from Two Datasets Alright, let's tackle this problem step by step. You've got two sets of column data, and you need to combine them to generate that new_col which pulls each tuple's element frequencies from the second dataset plus the tuple's own frequency. Here's a straightforward, efficient way to do this using pandas—perfect for this kind of tabular data work:
First, let's represent your raw data as pandas DataFrames (this makes manipulation way easier):
import pandas as pd # First dataset with tuple-freq pairs df1 = pd.DataFrame({ 'Col1': [('A','B'), ('C','C'), ('F','D')], 'Col2': [5, 4, 8] }) # Second dataset with element-freq pairs df2 = pd.DataFrame({ 'Col3': ['A', 'B', 'C', 'F', 'D'], 'Col4': [2, 5, 1, 2, 3] })
Next, let's convert the second dataset into a lookup dictionary—this lets us grab any element's frequency in constant time, no repeated searching required:
# Map each element to its frequency for quick lookups element_freq = df2.set_index('Col3')['Col4'].to_dict()
Now, we'll build a simple function to generate each entry in new_col, then apply it to every row in the first dataset:
def create_new_col(tuple_val, tuple_freq): # Get frequencies for both elements in the tuple first_elem_freq = element_freq[tuple_val[0]] second_elem_freq = element_freq[tuple_val[1]] # Format into the comma-separated string you need return f"{first_elem_freq},{second_elem_freq},{tuple_freq}" # Apply the function to each row and create the new column df1['new_col'] = df1.apply(lambda row: create_new_col(row['Col1'], row['Col2']), axis=1) # Trim down to just the columns specified in your desired output final_result = df1[['Col1', 'new_col']] print(final_result)
Running this code will output exactly what you're asking for:
Col1 new_col 0 ('A', 'B') 2,5,5 1 ('C', 'C') 1,1,4 2 ('F', 'D') 2,3,8
Quick breakdown of how this works:
- The dictionary
element_freqeliminates the need to scan the entire second dataset every time we need an element's frequency—this is way faster, especially if your data grows larger. - The
applymethod runs our custom logic on each row, pulling the right values and formatting them into the string format you specified. - We end up with a clean DataFrame that matches your requested output structure.
If you're not using pandas, you could also do this with basic Python loops and dictionaries, but pandas makes this kind of tabular manipulation far cleaner and more efficient.
内容的提问来源于stack exchange,提问作者user15649753

