Python Pandas:列表类型Noun列转新列累加点赞值报错求助
TypeError: unhashable type: 'list' for Pivoting List Columns in Pandas Hey there! That error makes total sense—pivot_table expects column values to be hashable (like strings or numbers), but lists are mutable and can’t be hashed. Let’s walk through how to fix this and get your desired output step by step.
Step 1: Break Down the Problem
Your Noun column holds lists of values, but we need to treat each individual noun as a separate entry paired with its row’s Thumbups value before we can pivot. That’s where pandas’ explode() method comes in handy.
Step 2: Explode the List Column
First, we’ll split each list in the Noun column into its own row, keeping the corresponding Thumbups value attached to each noun:
# First, recreate your sample DataFrame (for testing) import pandas as pd data = { "Review_Text": [ "Would be nice to be able to import files from ...", "No Offline Maps! It used to have offline maps ...", "Great application. Designed with very well tho...", "Great App. Nice and simple but accurate. Wish ...", "Save For Offline - This does not work. The rou...", "Since latest update app will not run. Subscrip...", "Great app. Love it! And all the things it does...", "I have paid for subscription but keeps telling...", "Error: The route cannot be save for no locatio...", "When try to restore my tracks it says \"unable ...", "Was a good app but since the update it only re..." ], "Noun": [ ["My", "Tracks", "app", "phone", "Google", "Drive", "import"], ["Offline", "Maps", "menu", "option", "video", "exchange"], ["application", "application"], ["Great", "App", "Nice", "Exported"], ["Save", "Offline", "route", "filesystem"], ["update", "app", "Subscription", "March", "application"], ["Great", "app", "Thank", "work"], ["subscription", "trial", "period"], ["Error", "route", "i", "GPS"], ["try", "file", "locally-1"], ["app", "update", "metre"] ], "Thumbups": [1.0, 18.0, 16.0, 0.0, 12.0, 9.0, 1.0, 0.0, 0.0, 0.0, 2.0] } latest_review = pd.DataFrame(data) # Explode the Noun column to split lists into individual rows exploded_df = latest_review.explode("Noun", ignore_index=True)
Step 3: Pivot to Sum Thumbups by Noun
Now that each noun is a single, hashable value in the Noun column, we can use pivot_table (or groupby + unstack) to calculate the total thumbups per noun:
Option 1: Using pivot_table
result = pd.pivot_table( exploded_df, columns="Noun", values="Thumbups", aggfunc="sum", fill_value=0.0 # Replace NaN with 0 for nouns not present in a row )
Option 2: Using groupby + unstack (Alternative)
If you prefer a more explicit grouping approach, this achieves the same result:
result = exploded_df.groupby("Noun")["Thumbups"].sum().unstack(fill_value=0.0)
Step 4: Merge Back to Original Data (Optional)
If you want the pivot columns added back to your original DataFrame (instead of just a summary), you can merge using a temporary index:
# Add a temporary index to the original DataFrame latest_review["temp_idx"] = latest_review.index # Explode while preserving the original index, then pivot exploded_with_idx = latest_review.explode("Noun", ignore_index=False).reset_index() pivot_with_idx = exploded_with_idx.pivot_table( index="temp_idx", columns="Noun", values="Thumbups", aggfunc="sum", fill_value=0.0 ) # Merge back to the original DataFrame and clean up final_df = latest_review.merge(pivot_with_idx, on="temp_idx").drop("temp_idx", axis=1)
Why This Works
explode()converts each list entry into a separate row, turning the unhashable list values into hashable strings thatpivot_tablecan work with.- The pivot operation then groups by each noun and sums the corresponding
Thumbupsvalues, exactly matching your requirement.
内容的提问来源于stack exchange,提问作者user2293224

