如何在PHP中拆分HTML为多段并保留标记(适配Google TTS配额)
str_truncate to Split HTML into All Valid Chunks Great question! That existing str_truncate function works perfectly for single truncations, but to split your HTML content into all valid, quota-compliant chunks (while preserving tag structure, whole words, and staying under your character limit), here's a practical, step-by-step approach:
1. Modify str_truncate to Return Critical Metadata
The original function only returns the truncated HTML. To iterate through chunks, we need it to return extra context to maintain tag validity and track progress:
- The list of open tags that were closed at the end of the truncation (so we can re-open them in the next chunk)
- The actual plain-text character count used in the truncated chunk (to stay within your quota)
- The unprocessed portion of the original HTML (stripped of the truncated content)
Here's a adjusted version of the function that returns this metadata:
function str_truncate_with_metadata($text, $length = 100, $ending = '', $exact = true, $considerHtml = false) { // Keep all original function logic here... // Instead of just returning $truncate, return an array with context return [ 'chunk' => $truncate, 'closed_tags' => $open_tags, // Tags that were closed to end the chunk 'used_plain_length' => $total_length - strlen($ending), 'remaining_html' => extract_remaining_html($text, $truncate, $open_tags) // Helper to get unprocessed content ]; } // Helper to extract the unprocessed HTML after truncation function extract_remaining_html($original, $truncated, $closed_tags) { // Reverse-engineer where the truncation stopped, accounting for closed tags $clean_truncated = preg_replace('/<\/' . implode('>|<\/', $closed_tags) . '>$/', '', $truncated); $remaining = substr($original, strpos($original, $clean_truncated) + strlen($clean_truncated)); return trim($remaining); }
2. Build an Iterative Chunking Wrapper
Create a main function that loops through the content, using the modified truncate function to generate each chunk, and re-establishes tag context for subsequent chunks.
Example wrapper:
function split_html_into_quota_chunks($html, $max_plain_length = 5000) { $chunks = []; $current_html = $html; $tag_prefix = ''; // Tracks tags to re-open for the next chunk while (true) { // Calculate plain-text length of remaining content $remaining_plain_length = strlen(preg_replace('/<.*?>/', '', $current_html)); if ($remaining_plain_length <= $max_plain_length) { // Add the final chunk with any necessary tag prefix $chunks[] = $tag_prefix . $current_html; break; } // Generate the next chunk $result = str_truncate_with_metadata( $tag_prefix . $current_html, $max_plain_length, '', // No ending needed since we're splitting, not truncating true, true ); $chunks[] = $result['chunk']; // Prepare tag prefix for next chunk: re-open tags that were closed $tag_prefix = ''; foreach (array_reverse($result['closed_tags']) as $tag) { $tag_prefix .= "<$tag>"; } // Update to unprocessed content $current_html = $result['remaining_html']; } return $chunks; }
3. Key Edge Cases to Handle
- Nested Tags: Ensure that if a chunk ends mid-nested tag (like
<div><span>...), the next chunk starts with<span>to maintain valid HTML. - Entity Characters: Keep the original function's logic for counting entities as single characters to avoid over-quota issues.
- Self-Closing Tags: Ignore tags like
<br/>or<img>when tracking open/closed tags, since they don't need to be re-opened. - Word Boundaries: Maintain the
$exact = truesetting to avoid splitting words across chunks.
4. Test with Your Example
Using your sample HTML and a 70-character plain-text limit, this setup would generate the 4 valid chunks you provided—each with complete HTML tags, whole words, and no broken structure.
内容的提问来源于stack exchange,提问作者alexander.khmelnitskiy

