You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Android中如何高效截断字符串至完整单词末尾(避免截断单词)

Hey there! Let's fix this string truncation issue so it reliably cuts off at the end of a complete word instead of mid-word, while keeping things efficient as requested.

First, let's break down why your current code isn't working

Your existing loop checks positions 140 to 150 for spaces, but it overwrites the truncation index every time it finds a space—meaning it ends up using the last space it finds (like 148 in your logs) instead of the first one after your base length. That's why you're getting mid-word cuts; the logic doesn't prioritize the closest valid word boundary. Also, you're reprocessing the HTML and string multiple times, which wastes efficiency, and you're not handling edge cases like when the string is shorter than your base length.

The Solution

Here's a streamlined, efficient approach that fixes these issues. It focuses on finding the first valid word boundary after your base length, avoids redundant string processing, and respects your requirement to avoid full-string traversals for long text:

int baseLength = 140;
int maxSearchOffset = 10; // Limit search to 10 chars after base length (matches your original 140-150 range)

// Preprocess the string ONCE to avoid redundant work
String rawContent = cureent.getTextContent().replaceAll("\n", "");
String processedStr = Html.fromHtml(rawContent).toString().trim();

int truncateIndex = baseLength;

// Handle case where string is already shorter than base length
if (processedStr.length() <= baseLength) {
    truncateIndex = processedStr.length();
} else {
    // Check if base length position is already a word boundary
    if (Character.isWhitespace(processedStr.charAt(baseLength))) {
        truncateIndex = baseLength;
    } else {
        // Search for the first space after base length (up to max offset)
        int endSearchAt = Math.min(baseLength + maxSearchOffset, processedStr.length() - 1);
        for (int i = baseLength; i <= endSearchAt; i++) {
            if (Character.isWhitespace(processedStr.charAt(i))) {
                truncateIndex = i;
                break; // Stop at the first valid boundary—no need to check further
            }
        }
        // If no space found in range, truncate at the end of the search range
        if (truncateIndex == baseLength) {
            truncateIndex = endSearchAt;
        }
    }
}

// Finalize the truncated string (handle special char and trim)
String truncatedDesc = processedStr.substring(0, truncateIndex)
        .replace((char) 65532, (char) 32)
        .trim();

item.setDescription(truncatedDesc);

Key Improvements

  • Single preprocessing pass: We parse HTML, replace newlines, and trim once instead of repeating the work multiple times—big win for efficiency.
  • Targeted search: We stop at the first space we find after the base length, ensuring we get the closest valid word boundary instead of the last one in the range.
  • Edge case handling: We account for strings shorter than the base length, and handle scenarios where no space exists in the search range (so we don't get stuck in an infinite loop or crash).
  • Efficient character checks: Using Character.isWhitespace() is faster than converting chars to strings for comparison, and works for all whitespace types (if you only want regular spaces, replace it with processedStr.charAt(i) == ' ').

Testing with your example

If we set baseLength = 14 for your sample string:

"And so, in honor of the antechinus’ tragic death, we give you a small sampling of the"

The code will check position 14 (which is the 'n' in "hon"), then search forward until it hits the space after "honor". It truncates at that space, giving you the desired result: "And so, in honor".

内容的提问来源于stack exchange,提问作者Demeteor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:17:11