You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何生成满足Subscribe_Year≤Last_Call的关联随机训练数据?

Got it, let's solve this constraint where Subscribe_Year is always ≤ Last_Call! Here are a few solid approaches, starting with the most efficient one:

1. Optimal Method: Shift the Negative Binomial Distribution

Since np.random.negative_binomial() generates non-negative integers, you can simply add your Subscribe_Year values to a set of non-negative negative binomial samples. This guarantees Last_Call >= Subscribe_Year without any loops or filtering—super clean and fast:

import pandas as pd
import numpy as np
import random

# First generate the Subscribe_Year column with your specified probabilities
sub_years = np.random.choice(range(0,4), 1932, p=[.2,.1,.3,.4], replace=True)

# Generate Last_Call as Subscribe_Year + a non-negative negative binomial value
# This ensures Last_Call is always >= Subscribe_Year
last_calls = sub_years + np.random.negative_binomial(5, 0.1, 1932)

# Assemble the final dataframe
Cust_Details = pd.DataFrame({
    "Cust_ID": random.sample(range(10000, 30000), k=1932),
    "Subscribe_Year": sub_years,
    "Last_Call": last_calls
})

Why this works: The negative binomial distribution here returns values ≥ 0, so adding Subscribe_Year ensures Last_Call is at least as large as the corresponding Subscribe_Year. No extra checks needed—this is the most efficient approach by far.

2. Per-Row Validation (For Strict Control)

If you want to keep the original negative binomial distribution shape but just filter out invalid values, you can generate Last_Call for each Subscribe_Year individually until you get a valid one:

import pandas as pd
import numpy as np
import random

# Generate Subscribe_Year first
sub_years = np.random.choice(range(0,4), 1932, p=[.2,.1,.3,.4], replace=True)

# Generate Last_Call values that meet the constraint
last_calls = []
for year in sub_years:
    # Keep generating until we get a Last_Call >= Subscribe_Year
    while True:
        call = np.random.negative_binomial(5, 0.1)
        if call >= year:
            last_calls.append(call)
            break

# Build the dataframe
Cust_Details = pd.DataFrame({
    "Cust_ID": random.sample(range(10000, 30000), k=1932),
    "Subscribe_Year": sub_years,
    "Last_Call": last_calls
})

Note: Since your negative binomial parameters give a mean of 45 (way larger than the max Subscribe_Year of 3), this loop will almost never run more than once—so it's still very efficient. This is useful if you need to preserve the exact distribution of Last_Call (minus the tiny fraction of values < 0-3).

3. Batch Filtering (For Simplicity)

Another option is to generate a batch of data, filter out rows that don't meet the constraint, and repeat until you hit your target sample size:

import pandas as pd
import numpy as np
import random

target_size = 1932
Cust_Details = pd.DataFrame()

while len(Cust_Details) < target_size:
    # Generate a batch of data (size = remaining samples needed)
    batch_size = target_size - len(Cust_Details)
    batch = pd.DataFrame({
        "Cust_ID": random.sample(range(10000, 30000), k=batch_size),
        "Subscribe_Year": np.random.choice(range(0,4), batch_size, p=[.2,.1,.3,.4], replace=True),
        "Last_Call": np.random.negative_binomial(5, 0.1, batch_size)
    })
    # Keep only valid rows where Subscribe_Year <= Last_Call
    valid_batch = batch[batch["Subscribe_Year"] <= batch["Last_Call"]]
    # Add valid rows to the dataframe
    Cust_Details = pd.concat([Cust_Details, valid_batch], ignore_index=True)

# Trim to exact target size (just in case we overshoot)
Cust_Details = Cust_Details.head(target_size)

This is simple to write and works well when most generated samples already meet the constraint (which they do here).

内容的提问来源于stack exchange,提问作者Brrrr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 19:52:35