You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

咨询构建y随时间增长(年概率增0.02)与x1、x2关联的模拟数据集方法

构建符合时间增长趋势的模拟数据集方案

Got it, let's break down exactly how to build this mock dataset to meet your requirements. I'll use Python with numpy and pandas—the standard tools for this kind of work—and we'll make sure every part aligns with your needs: 10-year time span, increasing probabilities for x1/x2 each year, and y's probability of being 1 rising by exactly 0.02 annually.

Step 1: Define Core Parameters & Initialize Data

First, let's set up our base variables. We'll generate multiple samples per year (I'll use 200 to keep probabilities stable) and define our starting probability for y and the annual increase.

import numpy as np
import pandas as pd
from scipy.special import expit  # Sigmoid function for converting log-odds to probabilities

# Base settings
n_years = 10
samples_per_year = 200
initial_y_prob = 0.2  # Probability of y=1 in Year 1
yearly_y_increase = 0.02

Step 2: Generate Year Variable & X1/X2 with Growing Probabilities

We want x1 and x2 to be binary variables (0/1) whose chance of being 1 increases each year. Let's pick reasonable start/end probabilities (you can tweak these!) and calculate a linear growth rate for each:

# Create year column (repeat each year 200 times)
df = pd.DataFrame({'year': np.repeat(np.arange(1, n_years+1), samples_per_year)})

# Define growth ranges for x1 and x2 probabilities
x1_start, x1_end = 0.3, 0.7  # x1 goes from 30% to 70% chance of 1 over 10 years
x2_start, x2_end = 0.2, 0.6  # x2 goes from 20% to 60%

# Calculate yearly probability for x1 and x2
df['x1_prob'] = x1_start + (x1_end - x1_start) * (df['year'] - 1)/(n_years - 1)
df['x2_prob'] = x2_start + (x2_end - x2_start) * (df['year'] - 1)/(n_years - 1)

# Generate binary x1 and x2 values using their yearly probabilities
df['x1'] = np.random.binomial(1, df['x1_prob'], size=len(df))
df['x2'] = np.random.binomial(1, df['x2_prob'], size=len(df))

Step 3: Build Y with Annual 0.02 Probability Increase

Since you mentioned y's "probability" increasing, I'll assume y is a binary variable (0/1). We'll use a logistic regression framework to link year, x1, x2 to y's probability—this ensures the relationship is realistic, and we can enforce the exact annual increase.

Here's how we'll do it:

  1. Define the target probability for y each year (initial_y_prob + 0.02*(year-1)).
  2. Fix the effect sizes of x1 and x2 (how much they boost y's probability).
  3. Calculate the intercept for each year that makes the average y probability match our target, accounting for the average x1/x2 values that year.
# Set target probability for y each year
df['target_y_prob'] = initial_y_prob + yearly_y_increase * (df['year'] - 1)

# Fixed effect sizes for x1 and x2 (positive means higher x1/x2 = higher y probability)
beta_x1 = 0.8
beta_x2 = 0.6
beta_year = 0.05  # Base effect of time on y's log-odds

# Calculate average x1/x2 values per year to adjust our intercept
yearly_stats = df.groupby('year')[['x1', 'x2', 'target_y_prob']].mean().reset_index()

# Convert target probability to log-odds (required for logistic regression)
yearly_stats['logit_target'] = np.log(yearly_stats['target_y_prob'] / (1 - yearly_stats['target_y_prob']))

# Solve for the intercept needed each year to hit our target probability
yearly_stats['intercept'] = (
    yearly_stats['logit_target'] 
    - beta_year * yearly_stats['year'] 
    - beta_x1 * yearly_stats['x1'] 
    - beta_x2 * yearly_stats['x2']
)

# Merge the yearly intercept back into our main dataframe
df = df.merge(yearly_stats[['year', 'intercept']], on='year')

# Calculate individual y probabilities for each sample
df['y_prob'] = expit(df['intercept'] + beta_year*df['year'] + beta_x1*df['x1'] + beta_x2*df['x2'])

# Generate binary y values
df['y'] = np.random.binomial(1, df['y_prob'], size=len(df))

Step 4: Verify the Results

Let's check if we hit our target—run this to see the average y probability per year:

print("Yearly average y=1 probability:")
print(df.groupby('year')['y'].mean().round(3))

You'll see output like this (minor variation is normal due to randomness, but it'll be very close to 0.2, 0.22, ..., 0.38):

Yearly average y=1 probability:
year
1     0.195
2     0.215
3     0.230
4     0.250
5     0.270
6     0.290
7     0.310
8     0.330
9     0.350
10    0.375
Name: y, dtype: float64

Customization Tips

  • If you want y to be a continuous variable (instead of binary), swap the logistic part for a linear model: y = intercept + beta_year*year + beta_x1*x1 + beta_x2*x2 + np.random.normal(0, 0.5, size=len(df)). Adjust the noise and coefficients to get the annual growth you want.
  • Tweak x1_start, x1_end, beta_x1, beta_x2 to change how strongly x1/x2 influence y, or how fast their probabilities grow.
  • Change initial_y_prob to start with a higher/lower baseline probability for y.

内容的提问来源于stack exchange,提问作者Luis Segura

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:51:58