如何用Keras与Pandas生成与原数组列分布偏差≤8%的重复数组
Got it, let's tackle this problem step by step. You need to generate 10x the original dataset (so 50 rows total, since your original has 5) while keeping each column's class distribution (for 0, 1, and None/NaN) within ±8% of the original. Here's a practical implementation using Pandas for data handling and Keras (optional, for model-ready batches) to generate the synthetic data, plus validation to ensure your constraints are met:
Step 1: Preprocess Original Data & Calculate Column Distributions
First, we'll convert your raw array into a Pandas DataFrame, then compute the exact distribution of 0, 1, and None (which becomes NaN in Pandas) for each column.
import pandas as pd import numpy as np import tensorflow as tf # Your original dataset original_data = [ [1, 1, 0, 0, None, 0, 1], [1, 0, 0, 0, None, 0, 1], [1, 1, None, 0, 1, 0, None], [1, 1, 1, 0, None, 0, 0], [1, 1, 0, None, 0, 0, 1] ] # Convert to DataFrame (None becomes NaN automatically) df = pd.DataFrame(original_data) # Calculate distribution of 0, 1, NaN for each column col_distributions = {} for col in df.columns: # Get counts including NaN counts = df[col].value_counts(dropna=False) total_rows = len(df[col]) # Convert counts to proportions dist = {val: count / total_rows for val, count in counts.items()} # Ensure all three classes (0,1,NaN) are present in the distribution dict for val in [0, 1, None]: if val not in dist: dist[val] = 0.0 col_distributions[col] = dist
Step 2: Define Allowed Distribution Ranges
We'll set up ±8% variance for each class's proportion in every column, making sure we don't go below 0% or above 100%.
allowed_variance = 0.08 # 8% col_ranges = {} for col, dist in col_distributions.items(): ranges = {} for val, prob in dist.items(): # Calculate min/max allowed proportion min_prob = max(0.0, prob - allowed_variance) max_prob = min(1.0, prob + allowed_variance) ranges[val] = (min_prob, max_prob) col_ranges[col] = ranges
Step 3: Generate Synthetic Data with Pandas
This function generates synthetic rows by sampling each column's values, adjusting the sampling probability within the allowed range, then normalizing to ensure valid probabilities.
def generate_synthetic_data(df, col_distributions, col_ranges, num_samples): synthetic_rows = [] for _ in range(num_samples): row = [] for col in df.columns: dist = col_distributions[col] ranges = col_ranges[col] # Adjust each class's probability within the allowed range adjusted_probs = {} for val, prob in dist.items(): min_p, max_p = ranges[val] adjusted_probs[val] = np.random.uniform(min_p, max_p) # Normalize probabilities to sum to 1 (critical for valid sampling) total_prob = sum(adjusted_probs.values()) normalized_probs = [adjusted_probs[val]/total_prob for val in [0,1,None]] # Sample the value for this column sampled_val = np.random.choice([0,1,None], p=normalized_probs) row.append(sampled_val) synthetic_rows.append(row) return pd.DataFrame(synthetic_rows) # Generate 10x the original data (5*10=50 rows) synthetic_df = generate_synthetic_data(df, col_distributions, col_ranges, num_samples=50)
Step 4: Optional - Keras Data Generator for Model Input
If you need to feed this synthetic data directly into a Keras model, use a Sequence class to create batches (and handle NaNs by replacing them with a placeholder like -1, since models can't process NaNs):
class SyntheticDataGenerator(tf.keras.utils.Sequence): def __init__(self, df, col_distributions, col_ranges, batch_size, num_samples): self.df = df self.col_distributions = col_distributions self.col_ranges = col_ranges self.batch_size = batch_size self.num_samples = num_samples self.indices = np.arange(num_samples) def __len__(self): # Number of batches per epoch return int(np.ceil(self.num_samples / self.batch_size)) def __getitem__(self, idx): # Get indices for current batch batch_indices = self.indices[idx*self.batch_size : (idx+1)*self.batch_size] batch_data = [] for _ in batch_indices: row = [] for col in self.df.columns: dist = self.col_distributions[col] ranges = col_ranges[col] adjusted_probs = {} for val, prob in dist.items(): min_p, max_p = ranges[val] adjusted_probs[val] = np.random.uniform(min_p, max_p) total_prob = sum(adjusted_probs.values()) normalized_probs = [adjusted_probs[val]/total_prob for val in [0,1,None]] sampled_val = np.random.choice([0,1,None], p=normalized_probs) row.append(sampled_val) batch_data.append(row) # Replace NaNs with -1 (or any placeholder your model accepts) batch_data = np.array(batch_data) batch_data = np.where(pd.isna(batch_data), -1, batch_data) return batch_data # Usage example: batch size 10, 50 total samples generator = SyntheticDataGenerator(df, col_distributions, col_ranges, batch_size=10, num_samples=50) first_batch = generator[0] # Get the first batch of 10 rows
Step 5: Validate the Synthetic Data
Finally, check that each column's distribution stays within the 8% variance limit:
def validate_distributions(original_df, synthetic_df, allowed_variance): for col in original_df.columns: # Get original proportions orig_props = original_df[col].value_counts(dropna=False, normalize=True) # Get synthetic proportions synth_props = synthetic_df[col].value_counts(dropna=False, normalize=True) print(f"\n--- Column {col} ---") for val in [0,1,None]: orig_p = orig_props.get(val, 0.0) synth_p = synth_props.get(val, 0.0) diff = abs(orig_p - synth_p) if diff > allowed_variance: print(f"⚠️ Value {val}: Original {orig_p:.2%}, Synthetic {synth_p:.2%}, Difference {diff:.2%} (EXCEEDS LIMIT)") else: print(f"✅ Value {val}: Original {orig_p:.2%}, Synthetic {synth_p:.2%}, Difference {diff:.2%} (WITHIN LIMIT)") # Run validation validate_distributions(df, synthetic_df, allowed_variance=0.08)
Key notes:
- We normalize the adjusted probabilities to ensure valid sampling (sum to 1)
- NaNs are handled by either keeping them in the DataFrame or replacing with a placeholder for Keras
- The validation step ensures your distribution constraint is respected
内容的提问来源于stack exchange,提问作者alextre

