如何通过Android Studio抓取空白HTML页面文本并在Android应用中展示
Hey there! I get it—sometimes those "blank" HTML pages that are just plain text wrapped in basic tags can trip up standard parsing approaches. Let’s walk through a straightforward, reliable solution to grab that text and show it in a TextView.
Step 1: Set Up Dependencies & Permissions
First, make sure you have the right tools to handle network requests and parsing. Add these dependencies to your module-level build.gradle (or build.gradle.kts):
// For lightweight, reliable network requests implementation "com.squareup.okhttp3:okhttp:4.11.0" // For safe, simple HTML parsing (avoids regex pitfalls) implementation "org.jsoup:jsoup:1.17.2"
Don’t forget to add the internet permission to your AndroidManifest.xml—without this, your app can’t reach the web page at all:
<uses-permission android:name="android.permission.INTERNET" />
Step 2: Fetch the Page Content in a Background Thread
Never make network requests on the main thread (it’ll cause app freezes or crashes!). Use Kotlin Coroutines (the modern, recommended approach) to handle this work in the background:
import okhttp3.OkHttpClient import okhttp3.Request import kotlinx.coroutines.CoroutineScope import kotlinx.coroutines.Dispatchers import kotlinx.coroutines.launch import kotlinx.coroutines.withContext // Call this method from your Activity/Fragment (e.g., in onCreate) fun fetchAndDisplayText() { val client = OkHttpClient() val request = Request.Builder() .url("https://pastebin.com/raw/Zuvfqdu9") .build() // Launch a coroutine on the IO dispatcher for network operations CoroutineScope(Dispatchers.IO).launch { try { val response = client.newCall(request).execute() val rawHtmlContent = response.body?.string() ?: "" // Parse the content and switch back to main thread to update UI val plainText = extractPlainText(rawHtmlContent) withContext(Dispatchers.Main) { yourTextViewId.text = plainText } } catch (e: Exception) { // Handle errors (e.g., no network, invalid response) e.printStackTrace() withContext(Dispatchers.Main) { yourTextViewId.text = "Failed to load content" } } } }
Step 3: Extract Plain Text from the Minimal HTML
The sample page wraps its text in a <pre> tag inside a basic HTML structure. You have two solid options here:
Option A: Jsoup (Most Reliable)
Jsoup is built for HTML parsing and handles edge cases way better than regex. Use it to target the <pre> tag directly:
import org.jsoup.Jsoup fun extractPlainText(rawHtml: String): String { val document = Jsoup.parse(rawHtml) // Grab the <pre> element that contains your target text val preElement = document.selectFirst("pre") // Return the text inside, or an empty string if the element isn't found return preElement?.text() ?: "" }
Option B: Simple String Replacement (Quick & Dependency-Free)
If you want to skip adding Jsoup, you can strip all HTML tags with a regex. Note: This works for simple pages but might break with more complex HTML:
fun extractPlainText(rawHtml: String): String { // Regex to remove all HTML tags and trim extra whitespace return rawHtml.replace(Regex("<[^>]+>"), "").trim() }
Why Your Previous Code Might Have Failed
- Main Thread Violations: If you tried making network calls on the main thread, Android would block the operation, leading to no content (or crashes).
- Overcomplicating Parsing: Heavy-duty HTML parsers might overprocess the minimal structure, missing the plain text entirely.
- Missing Permissions: Forgetting the
INTERNETpermission means your app can’t connect to the web page at all. - Null Handling Gaps: Failing to account for empty responses or missing elements would leave your
TextViewblank.
Final Notes
- Always handle errors (network issues, invalid responses) to give users clear feedback.
- For better performance, consider caching the response if you need to load the same content multiple times.
- If you’re using Java instead of Kotlin, replace Coroutines with an
ExecutorService(sinceAsyncTaskis deprecated).
内容的提问来源于stack exchange,提问作者Rooster Rooney

