左合并Dask DataFrame后结果为空数据帧的技术问题求助
Let's break down what's going wrong with your code and how to fix it step by step:
What's Causing the Empty Result?
The core issue is how you're handling the indexed DataFrame in Dask. When you use right_index=True for a merge, Dask expects the right DataFrame's index to be its partition key (meaning the data is partitioned based on that index).
In your code:
- You create
dask_pastwithnpartitions=4, but your originalpastDataFrame only has 3 rows. This splits the data into 4 lopsided partitions (3 with 1 row each, 1 completely empty). - Worse, the partitions aren't aligned with the
user_agentindex—Dask can't efficiently match rows across partitions when the index isn't the partition key, leading to an empty merge result.
Fix 1: Align Partitions with the Index
For indexed right-side merges, ensure the right Dask DataFrame is partitioned by its index. Since your past data is tiny, using a single partition makes the most sense (no need to overcomplicate with multiple partitions here):
import pandas as pd import dask.dataframe as dd # Original data setup raw_data = pd.DataFrame({ 'username': list('ab')*10, 'user_agent': list('cdef')*5, 'method': ['POST'] * 20, 'dst_port': [80]*20, 'dst': ['1.1.1.1']*20 }) past = pd.DataFrame({'user_agent': list('cde'), 'percent': [0.3, 0.3, 0.4]}).set_index('user_agent') # Create Dask DataFrames dask_raw = dd.from_pandas(raw_data, npartitions=4) # Key fix: Set index as the partition key (use 1 partition for small reference data) dask_past = dd.from_pandas(past, npartitions=1).set_index('user_agent') # Perform left merge merged_raw = dask_raw.merge(dask_past, how='left', left_on='user_agent', right_index=True) # View the computed result print(merged_raw.compute())
Fix 2: Merge on Columns Instead of Index (Simpler for Small Data)
If you don't strictly need to use the index for the merge, just keep user_agent as a column in both DataFrames. This avoids partition alignment issues entirely and is more straightforward for small datasets:
import pandas as pd import dask.dataframe as dd # Original data setup raw_data = pd.DataFrame({ 'username': list('ab')*10, 'user_agent': list('cdef')*5, 'method': ['POST'] * 20, 'dst_port': [80]*20, 'dst': ['1.1.1.1']*20 }) # Skip setting the index—keep user_agent as a regular column past = pd.DataFrame({'user_agent': list('cde'), 'percent': [0.3, 0.3, 0.4]}) # Create Dask DataFrames dask_raw = dd.from_pandas(raw_data, npartitions=4) dask_past = dd.from_pandas(past, npartitions=4) # Merge on the shared column merged_raw = dask_raw.merge(dask_past, how='left', on='user_agent') # View the computed result print(merged_raw.compute())
What to Expect in the Result
After fixing, you'll see rows with user_agent='f' have NaN for percent (since there's no matching entry in past), while rows with 'c', 'd', 'e' will have their corresponding percent values filled in correctly.
内容的提问来源于stack exchange,提问作者Apostolos

