如何在JavaScript中按权重从大数组抽取指定数量的数据
Got it, let's work through this problem efficiently—dealing with 300k entries means we need something way better than manual splitting. Here's a robust JavaScript solution that handles proportional sampling and gracefully fills gaps when a city doesn't have enough data, following your priority rules (same or higher weight cities first).
Step-by-Step Explanation
- Group & Shuffle First: We start by grouping data by city and shuffling each group once. This is efficient because we can just slice from the shuffled array later instead of re-sampling repeatedly.
- Calculate Quotas: Figure out how many entries we need from each city (12k for bauru/salvador, 9k for the others).
- Initial Sampling: Grab as many as we can from each city up to their quota.
- Fill Gaps: If any city can't meet its quota, sort cities with surplus data by weight (highest first) and pull extra entries from them until we hit our total of 60k.
- Final Shuffle: Mix all sampled entries together to avoid city-based clustering.
Full Code Implementation
// Configuration - tweak these as needed const TOTAL_SAMPLES = 60000; const CITY_QUOTA_PERCENTAGES = { bauru: 0.2, salvador: 0.2, belo: 0.15, natal: 0.15, belem: 0.15, santos: 0.15 }; const CITY_WEIGHTS = { bauru: 20, salvador: 20, belo: 15, natal: 15, belem: 15, santos: 15 }; // Helper: Group array items by city function groupByCity(data) { return data.reduce((groups, item) => { const city = item.city; if (!groups[city]) groups[city] = []; groups[city].push(item); return groups; }, {}); } // Helper: Fisher-Yates shuffle (in-place) for efficiency function shuffleArray(arr) { for (let i = arr.length - 1; i > 0; i--) { const j = Math.floor(Math.random() * (i + 1)); [arr[i], arr[j]] = [arr[j], arr[i]]; } return arr; } // Main sampling function function generateSample(data) { // 1. Group data by city and shuffle each group const cityGroups = groupByCity(data); Object.values(cityGroups).forEach(group => shuffleArray(group)); // 2. Calculate target quotas and track initial gaps const cityQuotas = {}; const initialGaps = {}; let totalGap = 0; Object.keys(CITY_QUOTA_PERCENTAGES).forEach(city => { cityQuotas[city] = Math.round(TOTAL_SAMPLES * CITY_QUOTA_PERCENTAGES[city]); const available = cityGroups[city]?.length || 0; initialGaps[city] = Math.max(0, cityQuotas[city] - available); totalGap += initialGaps[city]; }); // 3. Grab initial samples from each city const sampled = {}; Object.keys(cityQuotas).forEach(city => { const takeCount = Math.min(cityQuotas[city], cityGroups[city]?.length || 0); sampled[city] = cityGroups[city].slice(0, takeCount); // Update the group to remove sampled entries cityGroups[city] = cityGroups[city].slice(takeCount); }); // 4. Fill gaps using surplus cities (sorted by weight descending) if (totalGap > 0) { const surplusCities = Object.keys(cityGroups) .filter(city => cityGroups[city].length > 0) .sort((a, b) => CITY_WEIGHTS[b] - CITY_WEIGHTS[a]); let remainingGap = totalGap; for (const city of surplusCities) { if (remainingGap <= 0) break; const takeCount = Math.min(remainingGap, cityGroups[city].length); sampled[city] = sampled[city].concat(cityGroups[city].slice(0, takeCount)); cityGroups[city] = cityGroups[city].slice(takeCount); remainingGap -= takeCount; } // Fallback: If we still need more (edge case where total data is <60k) if (remainingGap > 0) { console.warn(`Warning: Still need ${remainingGap} samples. Falling back to random global sampling.`); const allRemaining = Object.values(cityGroups).flat(); shuffleArray(allRemaining); sampled.fallback = allRemaining.slice(0, remainingGap); } } // 5. Combine all samples and shuffle one final time const finalSample = Object.values(sampled).flat(); shuffleArray(finalSample); return finalSample; } // Usage example: // const sampledData = generateSample(your300kDataArray); // console.log(`Sampled ${sampledData.length} entries`);
Key Notes
- Efficiency: This runs in O(n) time (grouping is O(n), shuffling each group sums to O(n), and the rest are linear operations)—perfect for 300k entries.
- Quota Accuracy: Using
Math.roundmight lead to a total that's ±1 from 60k. If you need exact totals, adjust the quota calculation to distribute any remainder (e.g., add 1 to the first city with a fractional quota). - Priority Compliance: Surplus cities are sorted by weight first, so higher-weight cities (bauru/salvador) will fill gaps before same-weight ones, exactly as you requested.
- Randomness: Shuffling each group and the final sample ensures no bias in the order of entries.
内容的提问来源于stack exchange,提问作者Aks
相关产品推荐
相关产品推荐

