You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

开发多臂老虎机优化慈善捐赠:技术实现咨询

Hey there! Let's break down how to build this multi-armed bandit (MAB) system for optimizing charitable donation prompts—super impactful use case, by the way. Below’s a structured, practical guide to the technical implementation:

Core Components to Implement

1. Pick the Right MAB Algorithm for Continuous Rewards

Unlike standard click-through rate experiments (binary rewards), your system deals with continuous rewards (donation amounts from 0 to 100 cents). Here are the best algorithms to use:

Top Choices:

  • Thompson Sampling for Gaussian Rewards: Ideal for continuous, normally-distributed rewards. Each arm’s performance is modeled as a Gaussian distribution; you sample from each arm’s distribution and pick the one with the highest sampled value. Balances exploration (trying unproven prompts) and exploitation (sticking to high-performing ones) naturally.
  • UCB1-Tuned: A variant of Upper Confidence Bound that adapts to different reward variances between arms. Great if some prompts have wildly different donation patterns (e.g., one prompt gets lots of $1 donations, another gets many small 5-cent donations).
  • Contextual Bandits (LinUCB): Save this for later iterations—if you want to tailor prompts to user attributes (like past donation history, location), this algorithm lets you incorporate context to boost performance. Start with non-contextual MAB first to keep things simple.

Example Arm Tracking Class (Python)

You’ll need to track key stats for each prompt (A/B/C/D):

class DonationPromptArm:
    def __init__(self, prompt_id, prompt_text):
        self.prompt_id = prompt_id
        self.prompt_text = prompt_text
        self.total_donations = 0.0
        self.num_users = 0
        self.mean_donation = 0.0
        self.variance = 0.0  # For Gaussian Thompson Sampling

    def update(self, donation_amount):
        """Update stats with a new donation (amount in cents)"""
        self.num_users += 1
        self.total_donations += donation_amount
        
        # Update mean (Welford's algorithm for numerical stability)
        old_mean = self.mean_donation
        self.mean_donation = old_mean + (donation_amount - old_mean) / self.num_users
        
        # Update variance (only if we have >1 data point)
        if self.num_users > 1:
            self.variance = ((self.num_users - 2) * self.variance + 
                            (donation_amount - old_mean) * (donation_amount - self.mean_donation)) / (self.num_users - 1)

2. User Interaction & Data Pipeline

Your system needs a clear flow to serve prompts, collect donations, and update the MAB model in real time:

  • Serve Prompt: When a user arrives, the MAB algorithm selects the optimal prompt based on current stats.
  • Collect Donation: Show the user the $1 reward option and a way to input their donation (slider or input box restricted to 0-100 cents).
  • Update Model: Immediately send the donation amount to the backend to update the corresponding prompt’s stats. This ensures the next user gets the latest optimized choice.

Data Storage Tips:

  • Real-Time Stats: Use Redis to store each arm’s mean, variance, and user count—Redis is fast for read/write operations needed during prompt selection.
  • Historical Data: Use PostgreSQL to log every interaction (timestamp, prompt ID, donation amount) for later analysis (e.g., identifying trends over time).

3. System Architecture (Simple Version)

You don’t need a fancy setup to start—here’s a minimal stack:

Backend (Python FastAPI)

Create two core endpoints:

from fastapi import FastAPI
from pydantic import BaseModel
import redis
import json

app = FastAPI()
r = redis.Redis(host='localhost', port=6379, db=0)

# Initialize prompts (run once at startup)
prompts = [
    DonationPromptArm("A", "Text for prompt A..."),
    DonationPromptArm("B", "Text for prompt B..."),
    DonationPromptArm("C", "Text for prompt C..."),
    DonationPromptArm("D", "Text for prompt D...")
]
# Save initial stats to Redis
for p in prompts:
    r.set(p.prompt_id, json.dumps({
        "mean_donation": p.mean_donation,
        "variance": p.variance,
        "num_users": p.num_users
    }))

@app.get("/get-prompt")
def get_prompt():
    # Load current stats from Redis
    current_arms = []
    for p in prompts:
        stats = json.loads(r.get(p.prompt_id))
        arm = DonationPromptArm(p.prompt_id, p.prompt_text)
        arm.mean_donation = stats["mean_donation"]
        arm.variance = stats["variance"]
        arm.num_users = stats["num_users"]
        current_arms.append(arm)
    
    # Run Thompson Sampling to select arm
    ts = ThompsonSamplingGaussian(current_arms)
    selected_arm = ts.select_arm()
    return {"prompt_id": selected_arm.prompt_id, "prompt_text": selected_arm.prompt_text}

class DonationSubmission(BaseModel):
    prompt_id: str
    donation_amount: int  # Cents

@app.post("/submit-donation")
def submit_donation(submission: DonationSubmission):
    # Load current arm stats
    arm = next(p for p in prompts if p.prompt_id == submission.prompt_id)
    stats = json.loads(r.get(arm.prompt_id))
    arm.mean_donation = stats["mean_donation"]
    arm.variance = stats["variance"]
    arm.num_users = stats["num_users"]
    
    # Update arm with new donation
    arm.update(submission.donation_amount)
    
    # Save updated stats back to Redis
    r.set(arm.prompt_id, json.dumps({
        "mean_donation": arm.mean_donation,
        "variance": arm.variance,
        "num_users": arm.num_users
    }))
    
    # Log to PostgreSQL (omitted for brevity)
    return {"status": "success"}

Frontend

Build a simple HTML/JS page:

  • Fetch the prompt via /get-prompt on load.
  • Display the prompt and a donation input (slider from 0 to 100).
  • On submission, send the data to /submit-donation.

4. Monitor & Optimize the Experiment

  • Track Key Metrics: Use a dashboard (like Grafana) to monitor:
    • Average donation per prompt
    • Number of times each prompt is shown
    • Regret value (the difference between total donations if you’d always picked the best prompt, vs. actual total donations—lower = better algorithm performance)
  • Iterate: If a prompt consistently underperforms, replace its text with a new variant. If the algorithm isn’t exploring enough, adjust parameters (e.g., widen the prior distribution in Thompson Sampling).

5. Ethical & Compliance Checks

Don’t skip this—charitable work relies on trust:

  • Transparency: Tell users they’re part of an experiment to optimize donation prompts, and how their data is used.
  • Honesty: All prompts must be factually accurate—no misleading claims about where donations go.
  • Privacy: Follow regulations like GDPR/CCPA (anonymize user data where possible, let users opt out).
Example Thompson Sampling Implementation
import numpy as np
from scipy.stats import norm

class ThompsonSamplingGaussian:
    def __init__(self, arms):
        self.arms = arms

    def select_arm(self):
        sampled_rewards = []
        for arm in self.arms:
            if arm.num_users == 0:
                # Encourage exploration: sample from a wide prior (50 cents mean, 10 cents variance)
                sampled_reward = norm.rvs(loc=50, scale=10)
            else:
                # Sample from the arm's current Gaussian distribution
                sampled_reward = norm.rvs(loc=arm.mean_donation, scale=np.sqrt(arm.variance + 1e-6))  # Avoid division by zero
            sampled_rewards.append(sampled_reward)
        # Pick the arm with the highest sampled reward
        return self.arms[np.argmax(sampled_rewards)]

内容的提问来源于stack exchange,提问作者wwl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:49:55