Android:从HTML字符串提取指定标签内容并保留格式加载至View
问题:Jsoup提取HTML内容时如何保留
<p>标签格式加载到WebView 我获取到一个包含HTML内容的JSON长字符串,想要从中提取
<h1>和<p>标签内容:把<h1>内容赋值给Title TextView,把第一个<p>到指定div(示例里的<div class="m-top30 m-bottom20">)前的最后一个</p>的内容加载到WebView以保留段落排版。我已经用Jsoup实现了内容提取,但调用
.text()方法会把所有<p>内容合并成单行,没法保留<p>标签的段落格式,请问该怎么解决?
示例HTML代码
<div class="page-title-wrap sx-hide"> <div class="page-title clearfix"> <div class="col-lg-12"> <h1>Latest Deals</h1> </div> </div> </div> <div class="breadcrumb-wrapper"> <ul class="breadcrumb"> <li><a href="/Home">Home</a></li> <li><a href="/Deals">Deals</a></li> <li class="active">Great promotion! Is now RM 95 only! </li> </ul> </div> <div class="article outer clearfix"> <div class="col-sm-12"> <img alt="" title="Great promotion! Is now RM 95 only!" src=""> <h1>Great promotion! Is now RM 95 only!</h1> <p class="date">March 28th, 2017</p> <p><strong class="text-red"></strong></p> <p>This is the paragraht that shows the description of the promotion deals. You can write anything here. </p> <p>The buses offered by Alisan Golden Coach are in single deck or double deck. All of the buses are equip with air-conditioning and comfortable seats to ensure passengers are comfortable while travelling on the long journeys.</p> <p>Book your bus ticket before too late and enjoy the great saving.</p> <div class="m-top30 m-bottom20"> <a href="/home" class="btn btn-lg btn-orange">Home</a> </div> <div id="fb-root"></div> <script> (function(d, s, id) { var js, fjs = d.getElementsByTagName(s)[0]; if (d.getElementById(id)) return; js = d.createElement(s); js.id = id; js.async = true; js.src = ''; fjs.parentNode.insertBefore(js, fjs); }(document, 'script', 'facebook-jssdk'));</script> <div class="fb-share-button" data-href="http://google.com/" data-layout="button_count" data-size="large" data-mobile-iframe="true"> <a target="_blank" href="" class="fb-xfbml-parse-ignore">Share</a> </div> </div> </div>
当前使用的Jsoup代码
Document doc = Jsoup.parse(content); Elements eTitle = doc.getElementsByTag("h1"); Elements eBody = doc.getElementsByTag("p"); binding.fragmentWebview.loadData("<meta name=\"viewport\" content=\"width=device-width, initial-scale=1\">" + eBody.text(), "text/html; charset=utf-8","UTF-8");
解决方案
核心问题在于你用了.text()方法,它会把所有选中元素的文本内容提取出来并合并成纯文本,自然丢失了<p>标签带来的段落换行格式。我们需要保留<p>标签本身,同时只选取你需要的范围(第一个<p>到指定div前的最后一个<p>),具体步骤如下:
1. 提取正确的<h1>标题
注意示例里有两个<h1>,你需要的是文章内部的那个(也就是.article容器里的<h1>),所以可以更精准地定位:
// 定位到文章容器里的h1,避免拿到页面顶部的title Element articleContainer = doc.selectFirst("div.article.outer.clearfix"); String title = articleContainer.selectFirst("h1").text(); // 把title赋值给你的TextView binding.titleTextView.setText(title);
2. 提取指定范围的<p>标签并保留HTML格式
我们需要选中articleContainer里,从第一个<p>到<div class="m-top30 m-bottom20">之前的所有<p>元素,然后把这些元素的HTML代码拼接起来:
// 找到那个用来截断的目标div Element targetDiv = articleContainer.selectFirst("div.m-top30.m-bottom20"); // 选中目标div之前的所有p元素 Elements paragraphs = targetDiv.previousElementSiblings().select("p"); // 把这些p元素转换成完整的HTML字符串 StringBuilder bodyHtml = new StringBuilder(); for (Element p : paragraphs) { bodyHtml.append(p.outerHtml()); } // 加上viewport meta标签,适配移动端,然后加载到WebView String fullHtml = "<meta name=\"viewport\" content=\"width=device-width, initial-scale=1\">" + bodyHtml.toString(); binding.fragmentWebview.loadDataWithBaseURL(null, fullHtml, "text/html", "UTF-8", null);
为什么用outerHtml()而不是html()?
html()只会返回元素内部的内容,比如<p>xxx</p>会返回xxx;outerHtml()会返回元素本身加上内部内容,也就是完整的<p>xxx</p>标签,这样WebView就能识别并渲染出段落换行的效果。
额外优化:过滤空的<p>标签
示例里有个空的<p><strong class="text-red"></strong></p>,如果不想把这种空段落加载进去,可以在循环里加个判断:
for (Element p : paragraphs) { // 检查p标签是否有实际文本内容,避免空段落 if (!p.text().trim().isEmpty()) { bodyHtml.append(p.outerHtml()); } }
内容的提问来源于stack exchange,提问作者Shawn.Y
相关产品推荐
相关产品推荐

