开发多臂老虎机优化慈善捐赠:技术实现咨询
Hey there! Let's break down how to build this multi-armed bandit (MAB) system for optimizing charitable donation prompts—super impactful use case, by the way. Below’s a structured, practical guide to the technical implementation:
1. Pick the Right MAB Algorithm for Continuous Rewards
Unlike standard click-through rate experiments (binary rewards), your system deals with continuous rewards (donation amounts from 0 to 100 cents). Here are the best algorithms to use:
Top Choices:
- Thompson Sampling for Gaussian Rewards: Ideal for continuous, normally-distributed rewards. Each arm’s performance is modeled as a Gaussian distribution; you sample from each arm’s distribution and pick the one with the highest sampled value. Balances exploration (trying unproven prompts) and exploitation (sticking to high-performing ones) naturally.
- UCB1-Tuned: A variant of Upper Confidence Bound that adapts to different reward variances between arms. Great if some prompts have wildly different donation patterns (e.g., one prompt gets lots of $1 donations, another gets many small 5-cent donations).
- Contextual Bandits (LinUCB): Save this for later iterations—if you want to tailor prompts to user attributes (like past donation history, location), this algorithm lets you incorporate context to boost performance. Start with non-contextual MAB first to keep things simple.
Example Arm Tracking Class (Python)
You’ll need to track key stats for each prompt (A/B/C/D):
class DonationPromptArm: def __init__(self, prompt_id, prompt_text): self.prompt_id = prompt_id self.prompt_text = prompt_text self.total_donations = 0.0 self.num_users = 0 self.mean_donation = 0.0 self.variance = 0.0 # For Gaussian Thompson Sampling def update(self, donation_amount): """Update stats with a new donation (amount in cents)""" self.num_users += 1 self.total_donations += donation_amount # Update mean (Welford's algorithm for numerical stability) old_mean = self.mean_donation self.mean_donation = old_mean + (donation_amount - old_mean) / self.num_users # Update variance (only if we have >1 data point) if self.num_users > 1: self.variance = ((self.num_users - 2) * self.variance + (donation_amount - old_mean) * (donation_amount - self.mean_donation)) / (self.num_users - 1)
2. User Interaction & Data Pipeline
Your system needs a clear flow to serve prompts, collect donations, and update the MAB model in real time:
- Serve Prompt: When a user arrives, the MAB algorithm selects the optimal prompt based on current stats.
- Collect Donation: Show the user the $1 reward option and a way to input their donation (slider or input box restricted to 0-100 cents).
- Update Model: Immediately send the donation amount to the backend to update the corresponding prompt’s stats. This ensures the next user gets the latest optimized choice.
Data Storage Tips:
- Real-Time Stats: Use Redis to store each arm’s mean, variance, and user count—Redis is fast for read/write operations needed during prompt selection.
- Historical Data: Use PostgreSQL to log every interaction (timestamp, prompt ID, donation amount) for later analysis (e.g., identifying trends over time).
3. System Architecture (Simple Version)
You don’t need a fancy setup to start—here’s a minimal stack:
Backend (Python FastAPI)
Create two core endpoints:
from fastapi import FastAPI from pydantic import BaseModel import redis import json app = FastAPI() r = redis.Redis(host='localhost', port=6379, db=0) # Initialize prompts (run once at startup) prompts = [ DonationPromptArm("A", "Text for prompt A..."), DonationPromptArm("B", "Text for prompt B..."), DonationPromptArm("C", "Text for prompt C..."), DonationPromptArm("D", "Text for prompt D...") ] # Save initial stats to Redis for p in prompts: r.set(p.prompt_id, json.dumps({ "mean_donation": p.mean_donation, "variance": p.variance, "num_users": p.num_users })) @app.get("/get-prompt") def get_prompt(): # Load current stats from Redis current_arms = [] for p in prompts: stats = json.loads(r.get(p.prompt_id)) arm = DonationPromptArm(p.prompt_id, p.prompt_text) arm.mean_donation = stats["mean_donation"] arm.variance = stats["variance"] arm.num_users = stats["num_users"] current_arms.append(arm) # Run Thompson Sampling to select arm ts = ThompsonSamplingGaussian(current_arms) selected_arm = ts.select_arm() return {"prompt_id": selected_arm.prompt_id, "prompt_text": selected_arm.prompt_text} class DonationSubmission(BaseModel): prompt_id: str donation_amount: int # Cents @app.post("/submit-donation") def submit_donation(submission: DonationSubmission): # Load current arm stats arm = next(p for p in prompts if p.prompt_id == submission.prompt_id) stats = json.loads(r.get(arm.prompt_id)) arm.mean_donation = stats["mean_donation"] arm.variance = stats["variance"] arm.num_users = stats["num_users"] # Update arm with new donation arm.update(submission.donation_amount) # Save updated stats back to Redis r.set(arm.prompt_id, json.dumps({ "mean_donation": arm.mean_donation, "variance": arm.variance, "num_users": arm.num_users })) # Log to PostgreSQL (omitted for brevity) return {"status": "success"}
Frontend
Build a simple HTML/JS page:
- Fetch the prompt via
/get-prompton load. - Display the prompt and a donation input (slider from 0 to 100).
- On submission, send the data to
/submit-donation.
4. Monitor & Optimize the Experiment
- Track Key Metrics: Use a dashboard (like Grafana) to monitor:
- Average donation per prompt
- Number of times each prompt is shown
- Regret value (the difference between total donations if you’d always picked the best prompt, vs. actual total donations—lower = better algorithm performance)
- Iterate: If a prompt consistently underperforms, replace its text with a new variant. If the algorithm isn’t exploring enough, adjust parameters (e.g., widen the prior distribution in Thompson Sampling).
5. Ethical & Compliance Checks
Don’t skip this—charitable work relies on trust:
- Transparency: Tell users they’re part of an experiment to optimize donation prompts, and how their data is used.
- Honesty: All prompts must be factually accurate—no misleading claims about where donations go.
- Privacy: Follow regulations like GDPR/CCPA (anonymize user data where possible, let users opt out).
import numpy as np from scipy.stats import norm class ThompsonSamplingGaussian: def __init__(self, arms): self.arms = arms def select_arm(self): sampled_rewards = [] for arm in self.arms: if arm.num_users == 0: # Encourage exploration: sample from a wide prior (50 cents mean, 10 cents variance) sampled_reward = norm.rvs(loc=50, scale=10) else: # Sample from the arm's current Gaussian distribution sampled_reward = norm.rvs(loc=arm.mean_donation, scale=np.sqrt(arm.variance + 1e-6)) # Avoid division by zero sampled_rewards.append(sampled_reward) # Pick the arm with the highest sampled reward return self.arms[np.argmax(sampled_rewards)]
内容的提问来源于stack exchange,提问作者wwl

