Python Newspaper模块需求:提取带HTML标签的文本
Hey there! The a.text property in the Newspaper library intentionally strips all HTML tags to return clean plain text. To retrieve the article content with intact HTML tags (like <img>, <p>, headings, etc.), we need to use the library's underlying parsed DOM elements instead.
Modified Code Solution
Here's how to adjust your code to get the formatted HTML content:
from newspaper import Article from lxml import etree url = 'http://www.infomoney.com.br/mercados/acoes-e-indices/noticia/7345670/dow-jones-tem-nova-derrocada-puxa-ibovespa-para-segunda-semana' a = Article(url, language='pt') a.download() a.parse() # Convert the main article's DOM node to an HTML string article_html = etree.tostring(a.top_node, encoding='unicode', method='html') print(article_html)
Key Details:
a.top_node: This is the lxml Element object representing the main article content (the library filters out non-article content like headers/footers automatically).etree.tostring(): Converts the DOM node back to an HTML string, preserving all original tags and structure.
Customization Options:
If you want to target specific elements (like just <img> tags), you can use lxml's XPath queries:
# Extract all <img> tags from the main article img_elements = a.top_node.xpath('//img') for img in img_elements: print(etree.tostring(img, encoding='unicode'))
This approach gives you full control over the HTML content while leveraging Newspaper's built-in ability to identify the core article content.
内容的提问来源于stack exchange,提问作者Guilherme Hathy
相关产品推荐
相关产品推荐

