You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在JavaScript中按权重从大数组抽取指定数量的数据

Got it, let's work through this problem efficiently—dealing with 300k entries means we need something way better than manual splitting. Here's a robust JavaScript solution that handles proportional sampling and gracefully fills gaps when a city doesn't have enough data, following your priority rules (same or higher weight cities first).

Step-by-Step Explanation

  • Group & Shuffle First: We start by grouping data by city and shuffling each group once. This is efficient because we can just slice from the shuffled array later instead of re-sampling repeatedly.
  • Calculate Quotas: Figure out how many entries we need from each city (12k for bauru/salvador, 9k for the others).
  • Initial Sampling: Grab as many as we can from each city up to their quota.
  • Fill Gaps: If any city can't meet its quota, sort cities with surplus data by weight (highest first) and pull extra entries from them until we hit our total of 60k.
  • Final Shuffle: Mix all sampled entries together to avoid city-based clustering.

Full Code Implementation

// Configuration - tweak these as needed
const TOTAL_SAMPLES = 60000;
const CITY_QUOTA_PERCENTAGES = {
  bauru: 0.2,
  salvador: 0.2,
  belo: 0.15,
  natal: 0.15,
  belem: 0.15,
  santos: 0.15
};
const CITY_WEIGHTS = {
  bauru: 20,
  salvador: 20,
  belo: 15,
  natal: 15,
  belem: 15,
  santos: 15
};

// Helper: Group array items by city
function groupByCity(data) {
  return data.reduce((groups, item) => {
    const city = item.city;
    if (!groups[city]) groups[city] = [];
    groups[city].push(item);
    return groups;
  }, {});
}

// Helper: Fisher-Yates shuffle (in-place) for efficiency
function shuffleArray(arr) {
  for (let i = arr.length - 1; i > 0; i--) {
    const j = Math.floor(Math.random() * (i + 1));
    [arr[i], arr[j]] = [arr[j], arr[i]];
  }
  return arr;
}

// Main sampling function
function generateSample(data) {
  // 1. Group data by city and shuffle each group
  const cityGroups = groupByCity(data);
  Object.values(cityGroups).forEach(group => shuffleArray(group));

  // 2. Calculate target quotas and track initial gaps
  const cityQuotas = {};
  const initialGaps = {};
  let totalGap = 0;

  Object.keys(CITY_QUOTA_PERCENTAGES).forEach(city => {
    cityQuotas[city] = Math.round(TOTAL_SAMPLES * CITY_QUOTA_PERCENTAGES[city]);
    const available = cityGroups[city]?.length || 0;
    initialGaps[city] = Math.max(0, cityQuotas[city] - available);
    totalGap += initialGaps[city];
  });

  // 3. Grab initial samples from each city
  const sampled = {};
  Object.keys(cityQuotas).forEach(city => {
    const takeCount = Math.min(cityQuotas[city], cityGroups[city]?.length || 0);
    sampled[city] = cityGroups[city].slice(0, takeCount);
    // Update the group to remove sampled entries
    cityGroups[city] = cityGroups[city].slice(takeCount);
  });

  // 4. Fill gaps using surplus cities (sorted by weight descending)
  if (totalGap > 0) {
    const surplusCities = Object.keys(cityGroups)
      .filter(city => cityGroups[city].length > 0)
      .sort((a, b) => CITY_WEIGHTS[b] - CITY_WEIGHTS[a]);

    let remainingGap = totalGap;
    for (const city of surplusCities) {
      if (remainingGap <= 0) break;
      const takeCount = Math.min(remainingGap, cityGroups[city].length);
      sampled[city] = sampled[city].concat(cityGroups[city].slice(0, takeCount));
      cityGroups[city] = cityGroups[city].slice(takeCount);
      remainingGap -= takeCount;
    }

    // Fallback: If we still need more (edge case where total data is <60k)
    if (remainingGap > 0) {
      console.warn(`Warning: Still need ${remainingGap} samples. Falling back to random global sampling.`);
      const allRemaining = Object.values(cityGroups).flat();
      shuffleArray(allRemaining);
      sampled.fallback = allRemaining.slice(0, remainingGap);
    }
  }

  // 5. Combine all samples and shuffle one final time
  const finalSample = Object.values(sampled).flat();
  shuffleArray(finalSample);

  return finalSample;
}

// Usage example:
// const sampledData = generateSample(your300kDataArray);
// console.log(`Sampled ${sampledData.length} entries`);

Key Notes

  • Efficiency: This runs in O(n) time (grouping is O(n), shuffling each group sums to O(n), and the rest are linear operations)—perfect for 300k entries.
  • Quota Accuracy: Using Math.round might lead to a total that's ±1 from 60k. If you need exact totals, adjust the quota calculation to distribute any remainder (e.g., add 1 to the first city with a fractional quota).
  • Priority Compliance: Surplus cities are sorted by weight first, so higher-weight cities (bauru/salvador) will fill gaps before same-weight ones, exactly as you requested.
  • Randomness: Shuffling each group and the final sample ensures no bias in the order of entries.

内容的提问来源于stack exchange,提问作者Aks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 15:32:41