如何使用XPath提取HTML中article标签末尾的文本内容?
Let's break down why your existing XPath queries aren't working and walk through the correct approach to grab that target text.
Why Your Original XPaths Failed
- First XPath:
/article()[last()]
This has two critical issues:- The leading
/targets the root node of your HTML (which is<html>, not<article>), so it can't locate the article element at that level. article()is invalid syntax—you want to select the<article>element itself, not call a non-existent function namedarticle.
- The leading
- Second XPath:
//*[contains(@itemtype, 'http://schema.org/Article')]
This correctly selects the<article>element, but it returns the element itself rather than the specific text nodes you're targeting (the content that comes after<figure>and other child elements).
The Correct XPath to Grab Your Target Text
Your goal is to extract the direct text nodes inside the <article> that appear after its child elements (like <header>, <p class="lead">, <figure>). Here's the XPath to do exactly that:
//article[@itemtype='http://schema.org/Article']/text()[normalize-space() != '']
What This Does:
//article[@itemtype='http://schema.org/Article']: Precisely selects the<article>element with the exactitemtypevalue (more reliable thancontainssince the value is fixed)./text(): Fetches all direct text nodes inside the<article>—this ignores text nested inside child elements like<h1>or<p class="lead">.[normalize-space() != '']: Filters out empty or whitespace-only text nodes, which come from line breaks and indentation in the raw HTML.
If You Need a Single Combined String
If your tool/environment supports XPath 2.0+, you can join the text nodes into one clean string directly:
string-join(//article[@itemtype='http://schema.org/Article']/text()[normalize-space() != ''], ' ')
For XPath 1.0 environments (common in many scraping libraries), you'll need to fetch the individual text nodes and concatenate them in your code. Here's a quick Python example using lxml:
from lxml import etree # Your HTML content html = """ <div class="cbox"><article class="cf" itemscope itemtype="http://schema.org/Article"> <header> <h1 itemprop="headline">Anzündhilfen: So bringen Sie die Kohle zur Weissglut</h1> <em class="date"> <span class="my-color" itemprop="publisher">MyMagzine</span> 09/2018 vom <time datetime="2018-05-08" itemprop="datePublished">8. Mai 2018</time> | aktualisiert am <time datetime="2018-05-11" itemprop="dateModified">11. Mai 2018</time> </em> <p> von <span itemprop='author'>My Author</span> </p> </header> <p class="lead">Eine perfekte Glut ohne Rauch und Gestank bringen nur sogenannte Anzündkamine zustande. Aber zwei solche Produkte sind unsicher. </p> <figure class="image-box cf" itemscope itemtype="http://schema.org/ImageObject"> <img src="/image/?m=Artikel&rid=1113094&attr=bild&thumb=thumb_yRsBeq_resize_300_200.png" alt="Funken sprühen (Bild: CHRISTIAN BIRMELE)" itemprop="contentUrl"> <figcaption> <p itemprop="description">Funken sprühen (Bild: CHRISTIAN BIRMELE)</p> </figcaption> </figure> Here is my text that I would like to get out of this. Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla. <br /> <br />My Magazine has this title inbetween <br /> <br />Here is more text I also want to get our of this. [...]</p> </article> """ tree = etree.HTML(html) text_nodes = tree.xpath("//article[@itemtype='http://schema.org/Article']/text()[normalize-space() != '']") combined_text = " ".join([node.strip() for node in text_nodes]) print(combined_text)
This will output exactly the text you're targeting:
Here is my text that I would like to get out of this. Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla Bla. My Magazine has this title inbetween Here is more text I also want to get our of this. [...]
内容的提问来源于stack exchange,提问作者iKK

