基于词频的文本缩减算法及PHP实现方法技术咨询
Great question! Let's break down how to build a word frequency-based text reduction algorithm in PHP—one that preserves the relative frequency advantage of high-occurrence words, just like your example where house stays more frequent than book after reduction.
核心思路
The goal isn't just to cut down repeated words randomly; we need to maintain the "weight" of each word relative to others. For your sample input house house house house book book book, we shrink the most frequent word (house, 4 occurrences) to a target count (like 2), then scale other words proportionally (so book drops from 3 to 1) to keep house as the dominant term.
Step-by-Step Algorithm Logic
- Split text into words: Cleanly separate the input into individual words, handling extra spaces and optionally punctuation.
- Count word frequencies: Tally how many times each word appears in the text.
- Calculate scaling ratio: Determine how much to shrink frequencies based on your desired maximum occurrence count for the most frequent word.
- Set target occurrences: Apply the ratio to each word's count, making sure no word gets eliminated entirely (at least 1 occurrence).
- Reconstruct the reduced text: Either group repeated words (like your sample) or preserve the original scattered word order, depending on your needs.
PHP Implementation
Below are two flexible implementations—one for grouped word sequences and another for scattered word order.
Basic Version (Grouped Words, Matching Your Sample)
This outputs repeated words in blocks, exactly like your desired house house book result:
function reduceTextByFrequency($input, $targetMax = 2) { // Split text into words (handle multiple spaces and trim edges) $words = preg_split('/\s+/', trim($input)); if (empty($words)) return ''; // Count how many times each word appears $wordCounts = array_count_values($words); $maxCurrentCount = max($wordCounts); // No reduction needed if the most frequent word is already under the target if ($maxCurrentCount <= $targetMax) return $input; // Calculate how much to scale down all frequencies $scale = $targetMax / $maxCurrentCount; // Determine target counts (use floor to avoid over-scaling low-frequency words) $targetCounts = []; foreach ($wordCounts as $word => $count) { $targetCount = (int)floor($count * $scale); // Ensure every word appears at least once $targetCounts[$word] = max($targetCount, 1); } // Reconstruct text with grouped words, preserving the original word order $uniqueWordsInOrder = array_unique($words); $output = []; foreach ($uniqueWordsInOrder as $word) { $output = array_merge($output, array_fill(0, $targetCounts[$word], $word)); } return implode(' ', $output); } // Test your sample input $inputText = 'house house house house book book book'; echo reduceTextByFrequency($inputText); // Output: house house book
Advanced Version (Preserve Scattered Word Order)
If your input has words scattered (e.g., house book house book house book house), this version keeps the original sequence while stopping each word once it hits its target count:
function reduceTextPreserveOrder($input, $targetMax = 2) { $words = preg_split('/\s+/', trim($input)); if (empty($words)) return ''; $wordCounts = array_count_values($words); $maxCurrentCount = max($wordCounts); if ($maxCurrentCount <= $targetMax) return $input; $scale = $targetMax / $maxCurrentCount; $targetCounts = []; foreach ($wordCounts as $word => $count) { $targetCount = (int)floor($count * $scale); $targetCounts[$word] = max($targetCount, 1); } // Build output while tracking how many times each word has been added $output = []; $usedCounts = []; foreach ($words as $word) { if (!isset($usedCounts[$word])) $usedCounts[$word] = 0; if ($usedCounts[$word] < $targetCounts[$word]) { $output[] = $word; $usedCounts[$word]++; } } return implode(' ', $output); } // Test scattered input $scatteredInput = 'house book house book house book house'; echo reduceTextPreserveOrder($scatteredInput); // Output: house book house book
Key Customizations
- Adjust target maximum: Change the
$targetMaxparameter to control how many times the most frequent word appears (e.g., set to 3 forhouse house house book bookin your sample). - Handle punctuation: Modify the split regex to ignore punctuation, e.g.,
preg_split('/[^\w\']+/', trim($input))to keep apostrophes in words likedon't. - Case insensitivity: Add
$words = array_map('strtolower', $words);before counting if you wantHouseandhousetreated as the same word. - Rounding behavior: Replace
floor()withround()orceil()if you prefer different scaling for edge cases.
内容的提问来源于stack exchange,提问作者zeeks

