如何为Octoparse编写XPath以精准选取标题“Composição”下的<p>标签文本?
Solution for Extracting "Composição" Paragraph Text in Octoparse
Got it! If you need to extract just the text inside the <p> tag that sits right beneath the <h2> titled "Composição" using Octoparse, here's the perfect XPath expression to use:
//h2[normalize-space(.)='Composição']/following-sibling::p[1]/text()
Quick breakdown of how this works:
//h2[normalize-space(.)='Composição']: Finds the exact<h2>element by matching its normalized text (this handles any accidental extra spaces or line breaks in the heading)./following-sibling::p[1]: Targets the very first<p>tag that comes immediately after this<h2>—so you won’t accidentally grab other paragraphs further down the page./text(): Pulls only the raw text content inside that<p>tag, ignoring any underlying HTML structure.
Just paste this XPath into the selector field when setting up your data extraction rule in Octoparse, and it’ll fetch the composition text you need accurately.
内容的提问来源于stack exchange,提问作者over
相关产品推荐
相关产品推荐

