如何用Python的requests_html提取HTML<title>标签纯文本
如何用requests_html提取HTML标签的纯文本内容? <a class="header-anchor" href="#如何用requests-html提取html标签的纯文本内容?" aria-hidden="true">#</a></h1>
<p>你当前用<code>match[0].html</code>会返回元素及其后续兄弟节点的完整HTML,所以才会输出整个head部分。要获取<title>内的纯文本,有两种简单方法:</p>
<h2 id="方法一:使用element对象的text属性">方法一:使用Element对象的<code>text</code>属性 <a class="header-anchor" href="#方法一:使用element对象的text属性" aria-hidden="true">#</a></h2>
<p>直接提取匹配元素的文本内容,同时可以用<code>first=True</code>简化获取第一个匹配元素的代码:</p>
<pre class="hljs"><code class="language-python volc-pre-code">from requests_html import HTML
with open('simple.html') as html_file:
source = html_file.read()
html = HTML(html=source)
# 获取第一个title元素并提取其文本
title_element = html.find('title', first=True)
print(title_element.text)
</code></pre>
<h2 id="方法二:直接使用html对象的title属性">方法二:直接使用HTML对象的<code>title</code>属性 <a class="header-anchor" href="#方法二:直接使用html对象的title属性" aria-hidden="true">#</a></h2>
<p>requests_html的<code>HTML</code>对象内置了<code>title</code>属性,可直接返回<title>标签的纯文本,代码更简洁:</p>
<pre class="hljs"><code class="language-python volc-pre-code">from requests_html import HTML
with open('simple.html') as html_file:
source = html_file.read()
html = HTML(html=source)
print(html.title)
</code></pre>
<h3 id="附:你的html文件内容">附:你的HTML文件内容 <a class="header-anchor" href="#附:你的html文件内容" aria-hidden="true">#</a></h3>
<pre class="hljs"><code class="language-html volc-pre-code"><!doctype html>
<html class="no-js" lang="">
<head>
<title>Test - A Sample Website</title>
<!-- Other head elements -->
</head>
<body>
<!-- Body content -->
</body>
</html>
</code></pre>
<p>内容的提问来源于stack exchange,提问作者Terry</p>
相关产品推荐
相关产品推荐

