You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python爬取纽约时报时仅提取纯文本内容?

如何提取网页中的纯文本内容?

我写了一段爬取纽约时报文章的代码,但输出包含非文本信息,请问怎么调整才能只打印纯文本?

我的爬取代码:

url="https://www.nytimes.com/reuters/2018/04/05/business/05reuters-usa-tradechina-ip.html"
html = urlopen(url).read()
soup = BeautifulSoup(html,'html.parser')
def text_from_html(body):
    texts = soup.find_all('p',class_="story-body-text story-content")
    return texts
print(text_from_html(html)[0])

嘿,这问题很好解决!你现在返回的是BeautifulSoup的Tag对象集合,直接打印自然会带上HTML标签这类非文本内容。只要给每个Tag对象调用.get_text()方法(或者.text属性,两者效果相近),就能提取出里面的纯文本了。

给你调整好的代码,还加了些实用细节:

from urllib.request import urlopen
from bs4 import BeautifulSoup

url="https://www.nytimes.com/reuters/2018/04/05/business/05reuters-usa-tradechina-ip.html"
html = urlopen(url).read()
soup = BeautifulSoup(html,'html.parser')

def text_from_html(body):
    # 先定位所有目标段落标签
    text_tags = soup.find_all('p',class_="story-body-text story-content")
    # 逐个提取纯文本,strip=True去除多余的空格和换行
    return [tag.get_text(strip=True) for tag in text_tags]

# 打印第一段纯文本
print(text_from_html(html)[0])

# 如果想打印全部段落,可以用循环:
# for para in text_from_html(html):
#     print(para)
#     print("\n---\n")  # 加个分隔线,段落之间更清晰

简单说下细节:

  • get_text(strip=True)不仅能提取文本,还会自动清除文本首尾的空白字符和不必要的换行,让输出更整洁;要是你想保留原格式,去掉strip=True就行。
  • 别忘了导入需要的模块,不然代码会跑不起来哦!

内容的提问来源于stack exchange,提问作者Hello.World

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:15:58