如何截取HTML字符串以生成原始内容的预览版本?
Hey there! I’ve tackled similar HTML content trimming requirements before, so let’s walk through practical, actionable solutions for your preview needs. The core goal here is to extract raw text from the HTML (ignoring all tags), trim it to your 120-character limit, and optionally add an ellipsis for clarity.
Approach 1: Use Android’s Built-in HTML Utilities (Lightweight)
For most standard HTML generated by Android-RTEditor, the system’s Html class can handle text extraction reliably. Here’s how to implement it:
Step 1: Extract Plain Text from HTML
Use Html.fromHtml() to convert the HTML string into a Spanned object, then call toString() to get the raw text (all tags stripped out). Note the API version difference:
- For API 24+: Use
Html.fromHtml(html, Html.FROM_HTML_MODE_LEGACY) - For older versions: Use the deprecated
Html.fromHtml(html)
Step 2: Trim to 120 Characters
Once you have the plain text, check its length and truncate if needed. You can also clean up extra whitespace (like multiple newlines or spaces) for a cleaner preview.
Example Utility Method
fun generateHtmlPreview(htmlContent: String, maxLength: Int = 120): String { // Extract plain text from HTML val plainText = if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.N) { Html.fromHtml(htmlContent, Html.FROM_HTML_MODE_LEGACY).toString() } else { @Suppress("DEPRECATION") Html.fromHtml(htmlContent).toString() } // Clean up extra whitespace val cleanedText = plainText.replaceAll("\\s+", " ").trim() // Truncate and add ellipsis if needed return if (cleanedText.length <= maxLength) { cleanedText } else { cleanedText.substring(0, maxLength) + "..." } }
Approach 2: Use Jsoup for Complex HTML (More Reliable)
If your HTML has nested tags, special elements, or edge cases that the built-in Html class struggles with, Jsoup is a robust HTML parsing library that handles these scenarios seamlessly.
Step 1: Add Jsoup Dependency
In your app-level build.gradle (or build.gradle.kts):
dependencies { implementation 'org.jsoup:jsoup:1.17.2' }
Step 2: Extract and Trim Text
Jsoup’s text() method extracts all visible text from the HTML, ignoring tags and formatting. The rest of the truncation logic is the same as Approach 1.
Example Utility Method
fun generateHtmlPreviewWithJsoup(htmlContent: String, maxLength: Int = 120): String { // Parse HTML and extract plain text val plainText = Jsoup.parse(htmlContent).text() // Clean whitespace and truncate val cleanedText = plainText.replaceAll("\\s+", " ").trim() return if (cleanedText.length <= maxLength) { cleanedText } else { cleanedText.substring(0, maxLength) + "..." } }
Edge Cases to Consider
- Non-ASCII Characters: Both approaches handle Unicode characters (like emojis, accented letters) correctly, since they count as single characters in the string length.
- Empty or Whitespace-Only HTML: Add a check to return a default message (e.g., "No content") if the cleaned text is empty.
- Partial Words: If you want to avoid truncating mid-word, adjust the truncation logic to find the last space before the maxLength and cut there instead.
Bonus: Semi-Formatted Preview (If Needed)
If you want to retain basic formatting (like bold or italic) in the preview instead of stripping all tags, you can use Jsoup to traverse the HTML nodes, accumulate text until you hit the 120-character limit, and then build a truncated HTML snippet. This is more complex but useful if preview formatting matters.
内容的提问来源于stack exchange,提问作者android developer

