You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

左合并Dask DataFrame后结果为空数据帧的技术问题求助

Dask DataFrame Left Merge Returns Empty DataFrame - Fixes & Explanation

Let's break down what's going wrong with your code and how to fix it step by step:

What's Causing the Empty Result?

The core issue is how you're handling the indexed DataFrame in Dask. When you use right_index=True for a merge, Dask expects the right DataFrame's index to be its partition key (meaning the data is partitioned based on that index).

In your code:

  • You create dask_past with npartitions=4, but your original past DataFrame only has 3 rows. This splits the data into 4 lopsided partitions (3 with 1 row each, 1 completely empty).
  • Worse, the partitions aren't aligned with the user_agent index—Dask can't efficiently match rows across partitions when the index isn't the partition key, leading to an empty merge result.

Fix 1: Align Partitions with the Index

For indexed right-side merges, ensure the right Dask DataFrame is partitioned by its index. Since your past data is tiny, using a single partition makes the most sense (no need to overcomplicate with multiple partitions here):

import pandas as pd
import dask.dataframe as dd

# Original data setup
raw_data = pd.DataFrame({
    'username': list('ab')*10, 
    'user_agent': list('cdef')*5, 
    'method': ['POST'] * 20, 
    'dst_port': [80]*20, 
    'dst': ['1.1.1.1']*20
})
past = pd.DataFrame({'user_agent': list('cde'), 'percent': [0.3, 0.3, 0.4]}).set_index('user_agent')

# Create Dask DataFrames
dask_raw = dd.from_pandas(raw_data, npartitions=4)
# Key fix: Set index as the partition key (use 1 partition for small reference data)
dask_past = dd.from_pandas(past, npartitions=1).set_index('user_agent')

# Perform left merge
merged_raw = dask_raw.merge(dask_past, how='left', left_on='user_agent', right_index=True)

# View the computed result
print(merged_raw.compute())

Fix 2: Merge on Columns Instead of Index (Simpler for Small Data)

If you don't strictly need to use the index for the merge, just keep user_agent as a column in both DataFrames. This avoids partition alignment issues entirely and is more straightforward for small datasets:

import pandas as pd
import dask.dataframe as dd

# Original data setup
raw_data = pd.DataFrame({
    'username': list('ab')*10, 
    'user_agent': list('cdef')*5, 
    'method': ['POST'] * 20, 
    'dst_port': [80]*20, 
    'dst': ['1.1.1.1']*20
})
# Skip setting the index—keep user_agent as a regular column
past = pd.DataFrame({'user_agent': list('cde'), 'percent': [0.3, 0.3, 0.4]})

# Create Dask DataFrames
dask_raw = dd.from_pandas(raw_data, npartitions=4)
dask_past = dd.from_pandas(past, npartitions=4)

# Merge on the shared column
merged_raw = dask_raw.merge(dask_past, how='left', on='user_agent')

# View the computed result
print(merged_raw.compute())

What to Expect in the Result

After fixing, you'll see rows with user_agent='f' have NaN for percent (since there's no matching entry in past), while rows with 'c', 'd', 'e' will have their corresponding percent values filled in correctly.

内容的提问来源于stack exchange,提问作者Apostolos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:40:17