如何生成满足Subscribe_Year≤Last_Call的关联随机训练数据?
Got it, let's solve this constraint where Subscribe_Year is always ≤ Last_Call! Here are a few solid approaches, starting with the most efficient one:
1. Optimal Method: Shift the Negative Binomial Distribution
Since np.random.negative_binomial() generates non-negative integers, you can simply add your Subscribe_Year values to a set of non-negative negative binomial samples. This guarantees Last_Call >= Subscribe_Year without any loops or filtering—super clean and fast:
import pandas as pd import numpy as np import random # First generate the Subscribe_Year column with your specified probabilities sub_years = np.random.choice(range(0,4), 1932, p=[.2,.1,.3,.4], replace=True) # Generate Last_Call as Subscribe_Year + a non-negative negative binomial value # This ensures Last_Call is always >= Subscribe_Year last_calls = sub_years + np.random.negative_binomial(5, 0.1, 1932) # Assemble the final dataframe Cust_Details = pd.DataFrame({ "Cust_ID": random.sample(range(10000, 30000), k=1932), "Subscribe_Year": sub_years, "Last_Call": last_calls })
Why this works: The negative binomial distribution here returns values ≥ 0, so adding Subscribe_Year ensures Last_Call is at least as large as the corresponding Subscribe_Year. No extra checks needed—this is the most efficient approach by far.
2. Per-Row Validation (For Strict Control)
If you want to keep the original negative binomial distribution shape but just filter out invalid values, you can generate Last_Call for each Subscribe_Year individually until you get a valid one:
import pandas as pd import numpy as np import random # Generate Subscribe_Year first sub_years = np.random.choice(range(0,4), 1932, p=[.2,.1,.3,.4], replace=True) # Generate Last_Call values that meet the constraint last_calls = [] for year in sub_years: # Keep generating until we get a Last_Call >= Subscribe_Year while True: call = np.random.negative_binomial(5, 0.1) if call >= year: last_calls.append(call) break # Build the dataframe Cust_Details = pd.DataFrame({ "Cust_ID": random.sample(range(10000, 30000), k=1932), "Subscribe_Year": sub_years, "Last_Call": last_calls })
Note: Since your negative binomial parameters give a mean of 45 (way larger than the max Subscribe_Year of 3), this loop will almost never run more than once—so it's still very efficient. This is useful if you need to preserve the exact distribution of Last_Call (minus the tiny fraction of values < 0-3).
3. Batch Filtering (For Simplicity)
Another option is to generate a batch of data, filter out rows that don't meet the constraint, and repeat until you hit your target sample size:
import pandas as pd import numpy as np import random target_size = 1932 Cust_Details = pd.DataFrame() while len(Cust_Details) < target_size: # Generate a batch of data (size = remaining samples needed) batch_size = target_size - len(Cust_Details) batch = pd.DataFrame({ "Cust_ID": random.sample(range(10000, 30000), k=batch_size), "Subscribe_Year": np.random.choice(range(0,4), batch_size, p=[.2,.1,.3,.4], replace=True), "Last_Call": np.random.negative_binomial(5, 0.1, batch_size) }) # Keep only valid rows where Subscribe_Year <= Last_Call valid_batch = batch[batch["Subscribe_Year"] <= batch["Last_Call"]] # Add valid rows to the dataframe Cust_Details = pd.concat([Cust_Details, valid_batch], ignore_index=True) # Trim to exact target size (just in case we overshoot) Cust_Details = Cust_Details.head(target_size)
This is simple to write and works well when most generated samples already meet the constraint (which they do here).
内容的提问来源于stack exchange,提问作者Brrrr

