如何用R语言获取网页<head>中的<title>标签内容
R提取网页内容的正确方法 <a class="header-anchor" href="#r提取网页内容的正确方法" aria-hidden="true">#</a></h1>
<ul>
<li>
<p>你用HEAD方法失败的原因很直接:HEAD请求<strong>仅返回响应头信息</strong>,不会获取网页的HTML正文内容,而<title>标签是嵌在HTML正文的<head>里的,自然拿不到。</p>
</li>
<li>
<p>正确的思路是用GET请求获取完整的HTML页面,再通过HTML解析工具提取<title>文本。下面是两种常用实现方式:</p>
</li>
</ul>
<h2 id="方式一:使用httr-xml2">方式一:使用httr + xml2 <a class="header-anchor" href="#方式一:使用httr-xml2" aria-hidden="true">#</a></h2>
<pre class="hljs"><code class="language-r volc-pre-code">library(httr)
library(xml2)
# 发送GET请求获取页面内容
response <- GET("目标网页URL")
# 解析HTML内容
html_content <- read_html(content(response, "text"))
# 提取title文本
page_title <- xml_text(xml_find_first(html_content, "//title"))
print(page_title)
</code></pre>
<h2 id="方式二:使用httr-rvest(更简洁)">方式二:使用httr + rvest(更简洁) <a class="header-anchor" href="#方式二:使用httr-rvest(更简洁)" aria-hidden="true">#</a></h2>
<pre class="hljs"><code class="language-r volc-pre-code">library(httr)
library(rvest)
response <- GET("目标网页URL")
page_title <- response %>%
content("text") %>%
read_html() %>%
html_element("title") %>%
html_text()
print(page_title)
</code></pre>
<ul>
<li>注意事项:
<ul>
<li>部分网站会反爬,可在GET请求中添加<code>user_agent()</code>参数模拟浏览器,比如<code>GET(url, user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"))</code></li>
<li>若遇编码问题,可在<code>content()</code>里指定编码,比如<code>content(response, "text", encoding = "UTF-8")</code></li>
</ul>
</li>
</ul>
<p>内容的提问来源于stack exchange,提问作者broti</p>
相关产品推荐
相关产品推荐

