You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Newspaper模块需求:提取带HTML标签的文本

Extract HTML Content (Including <img> Tags) with Python Newspaper

Hey there! The a.text property in the Newspaper library intentionally strips all HTML tags to return clean plain text. To retrieve the article content with intact HTML tags (like <img>, <p>, headings, etc.), we need to use the library's underlying parsed DOM elements instead.

Modified Code Solution

Here's how to adjust your code to get the formatted HTML content:

from newspaper import Article
from lxml import etree

url = 'http://www.infomoney.com.br/mercados/acoes-e-indices/noticia/7345670/dow-jones-tem-nova-derrocada-puxa-ibovespa-para-segunda-semana'
a = Article(url, language='pt')
a.download()
a.parse()

# Convert the main article's DOM node to an HTML string
article_html = etree.tostring(a.top_node, encoding='unicode', method='html')
print(article_html)

Key Details:

  • a.top_node: This is the lxml Element object representing the main article content (the library filters out non-article content like headers/footers automatically).
  • etree.tostring(): Converts the DOM node back to an HTML string, preserving all original tags and structure.

Customization Options:

If you want to target specific elements (like just <img> tags), you can use lxml's XPath queries:

# Extract all <img> tags from the main article
img_elements = a.top_node.xpath('//img')
for img in img_elements:
    print(etree.tostring(img, encoding='unicode'))

This approach gives you full control over the HTML content while leveraging Newspaper's built-in ability to identify the core article content.

内容的提问来源于stack exchange,提问作者Guilherme Hathy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:58:50