You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大数据量下统计同一订单中物料编号对的出现次数

Absolutely, this is totally achievable—and it’s actually a classic use case for co-occurrence analysis! Your dataset size (30,000 orders, 700 unique items) is well within the range of standard data processing tools, so you won’t run into performance issues with the 700x700 matrix end result.

How to Pull This Off

Let’s walk through the core steps, plus a practical code example to make it concrete:

1. Prep Your Order-Item Relationships

First, you’ll want to group your data by order number to get a list of all items associated with each individual order. This makes it easy to generate item pairs per order.

2. Generate Item Pairs

For each order’s item list, create all unique unordered pairs (since "Item 10 & Item 11" is the same as "Item 11 & Item 10" for your counting needs). Using unordered pairs avoids double-counting the same combination in reverse.

3. Count Co-Occurrences & Build the Matrix

Once you have all pairs across every order, count how many times each pair appears. Then pivot this count data into your desired matrix format, where rows and columns are item IDs, and cell values are co-occurrence counts.

Practical Example (Python + Pandas)

Here’s a quick implementation using pandas—super straightforward for your dataset size:

import pandas as pd
from itertools import combinations

# Load your actual data instead of this sample
df = pd.DataFrame({
    '订单号': ['OD001', 'OD001', 'OD002', 'OD002', 'OD002', 'OD003'],
    '物料编号': ['10', '11', '10', '11', '12', '10']
})

# Group items by order number
order_item_groups = df.groupby('订单号')['物料编号'].apply(list)

# Generate all unordered item pairs per order
co_occur_pairs = []
for items in order_item_groups:
    # Sort items first to ensure consistent pair order (e.g., ('10','11') not ('11','10'))
    sorted_items = sorted(items)
    pairs = combinations(sorted_items, 2)
    co_occur_pairs.extend(pairs)

# Count pair occurrences
pair_counts = pd.DataFrame(co_occur_pairs, columns=['物料A', '物料B']).value_counts().reset_index(name='共现次数')

# Convert to matrix format
co_occur_matrix = pair_counts.pivot(index='物料A', columns='物料B', values='共现次数').fillna(0)

# Optional: Make the matrix symmetric (fill upper triangle with lower triangle values)
co_occur_matrix = co_occur_matrix + co_occur_matrix.T
Quick Notes
  • If your item IDs are numeric, convert them to strings first to avoid weird sorting or indexing issues.
  • The 700x700 matrix will have 490,000 cells—this is trivial for pandas to handle, even on a regular laptop.
  • If you need to include "self-co-occurrences" (same item appearing multiple times in one order), you can adjust the code to use combinations_with_replacement instead of combinations.

内容的提问来源于stack exchange,提问作者M.A

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:40:45