You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Keras与Pandas生成与原数组列分布偏差≤8%的重复数组

Got it, let's tackle this problem step by step. You need to generate 10x the original dataset (so 50 rows total, since your original has 5) while keeping each column's class distribution (for 0, 1, and None/NaN) within ±8% of the original. Here's a practical implementation using Pandas for data handling and Keras (optional, for model-ready batches) to generate the synthetic data, plus validation to ensure your constraints are met:

Step 1: Preprocess Original Data & Calculate Column Distributions

First, we'll convert your raw array into a Pandas DataFrame, then compute the exact distribution of 0, 1, and None (which becomes NaN in Pandas) for each column.

import pandas as pd
import numpy as np
import tensorflow as tf

# Your original dataset
original_data = [
    [1, 1, 0, 0, None, 0, 1],
    [1, 0, 0, 0, None, 0, 1],
    [1, 1, None, 0, 1, 0, None],
    [1, 1, 1, 0, None, 0, 0],
    [1, 1, 0, None, 0, 0, 1]
]

# Convert to DataFrame (None becomes NaN automatically)
df = pd.DataFrame(original_data)

# Calculate distribution of 0, 1, NaN for each column
col_distributions = {}
for col in df.columns:
    # Get counts including NaN
    counts = df[col].value_counts(dropna=False)
    total_rows = len(df[col])
    # Convert counts to proportions
    dist = {val: count / total_rows for val, count in counts.items()}
    # Ensure all three classes (0,1,NaN) are present in the distribution dict
    for val in [0, 1, None]:
        if val not in dist:
            dist[val] = 0.0
    col_distributions[col] = dist

Step 2: Define Allowed Distribution Ranges

We'll set up ±8% variance for each class's proportion in every column, making sure we don't go below 0% or above 100%.

allowed_variance = 0.08  # 8%
col_ranges = {}

for col, dist in col_distributions.items():
    ranges = {}
    for val, prob in dist.items():
        # Calculate min/max allowed proportion
        min_prob = max(0.0, prob - allowed_variance)
        max_prob = min(1.0, prob + allowed_variance)
        ranges[val] = (min_prob, max_prob)
    col_ranges[col] = ranges

Step 3: Generate Synthetic Data with Pandas

This function generates synthetic rows by sampling each column's values, adjusting the sampling probability within the allowed range, then normalizing to ensure valid probabilities.

def generate_synthetic_data(df, col_distributions, col_ranges, num_samples):
    synthetic_rows = []
    for _ in range(num_samples):
        row = []
        for col in df.columns:
            dist = col_distributions[col]
            ranges = col_ranges[col]
            
            # Adjust each class's probability within the allowed range
            adjusted_probs = {}
            for val, prob in dist.items():
                min_p, max_p = ranges[val]
                adjusted_probs[val] = np.random.uniform(min_p, max_p)
            
            # Normalize probabilities to sum to 1 (critical for valid sampling)
            total_prob = sum(adjusted_probs.values())
            normalized_probs = [adjusted_probs[val]/total_prob for val in [0,1,None]]
            
            # Sample the value for this column
            sampled_val = np.random.choice([0,1,None], p=normalized_probs)
            row.append(sampled_val)
        synthetic_rows.append(row)
    
    return pd.DataFrame(synthetic_rows)

# Generate 10x the original data (5*10=50 rows)
synthetic_df = generate_synthetic_data(df, col_distributions, col_ranges, num_samples=50)

Step 4: Optional - Keras Data Generator for Model Input

If you need to feed this synthetic data directly into a Keras model, use a Sequence class to create batches (and handle NaNs by replacing them with a placeholder like -1, since models can't process NaNs):

class SyntheticDataGenerator(tf.keras.utils.Sequence):
    def __init__(self, df, col_distributions, col_ranges, batch_size, num_samples):
        self.df = df
        self.col_distributions = col_distributions
        self.col_ranges = col_ranges
        self.batch_size = batch_size
        self.num_samples = num_samples
        self.indices = np.arange(num_samples)
    
    def __len__(self):
        # Number of batches per epoch
        return int(np.ceil(self.num_samples / self.batch_size))
    
    def __getitem__(self, idx):
        # Get indices for current batch
        batch_indices = self.indices[idx*self.batch_size : (idx+1)*self.batch_size]
        batch_data = []
        
        for _ in batch_indices:
            row = []
            for col in self.df.columns:
                dist = self.col_distributions[col]
                ranges = col_ranges[col]
                
                adjusted_probs = {}
                for val, prob in dist.items():
                    min_p, max_p = ranges[val]
                    adjusted_probs[val] = np.random.uniform(min_p, max_p)
                
                total_prob = sum(adjusted_probs.values())
                normalized_probs = [adjusted_probs[val]/total_prob for val in [0,1,None]]
                sampled_val = np.random.choice([0,1,None], p=normalized_probs)
                row.append(sampled_val)
            
            batch_data.append(row)
        
        # Replace NaNs with -1 (or any placeholder your model accepts)
        batch_data = np.array(batch_data)
        batch_data = np.where(pd.isna(batch_data), -1, batch_data)
        
        return batch_data

# Usage example: batch size 10, 50 total samples
generator = SyntheticDataGenerator(df, col_distributions, col_ranges, batch_size=10, num_samples=50)
first_batch = generator[0]  # Get the first batch of 10 rows

Step 5: Validate the Synthetic Data

Finally, check that each column's distribution stays within the 8% variance limit:

def validate_distributions(original_df, synthetic_df, allowed_variance):
    for col in original_df.columns:
        # Get original proportions
        orig_props = original_df[col].value_counts(dropna=False, normalize=True)
        # Get synthetic proportions
        synth_props = synthetic_df[col].value_counts(dropna=False, normalize=True)
        
        print(f"\n--- Column {col} ---")
        for val in [0,1,None]:
            orig_p = orig_props.get(val, 0.0)
            synth_p = synth_props.get(val, 0.0)
            diff = abs(orig_p - synth_p)
            
            if diff > allowed_variance:
                print(f"⚠️  Value {val}: Original {orig_p:.2%}, Synthetic {synth_p:.2%}, Difference {diff:.2%} (EXCEEDS LIMIT)")
            else:
                print(f"✅ Value {val}: Original {orig_p:.2%}, Synthetic {synth_p:.2%}, Difference {diff:.2%} (WITHIN LIMIT)")

# Run validation
validate_distributions(df, synthetic_df, allowed_variance=0.08)

Key notes:

  • We normalize the adjusted probabilities to ensure valid sampling (sum to 1)
  • NaNs are handled by either keeping them in the DataFrame or replacing with a placeholder for Keras
  • The validation step ensures your distribution constraint is respected

内容的提问来源于stack exchange,提问作者alextre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:14:16