请求编写用于提取指定HTML中总页数(目标结果为111)的XPath表达式
XPath Expression to Extract Total Pages
Here's an XPath expression that will reliably pull out the total page number (111) from your provided HTML snippet:
normalize-space(substring-after(substring-before(//span[@id='ctl00_ContentPlaceHolder1_PagerControl1_litRecords']/text(), ']'), 'Total pages: '))
Breakdown of how it works:
- Target the span element:
//span[@id='ctl00_ContentPlaceHolder1_PagerControl1_litRecords']uses the span's unique ID to pinpoint it accurately—this is the most reliable way to locate the element since IDs are intended to be unique. - Retrieve the full text content:
/text()grabs all the text inside the span:"Showing 1 - 500 of 55385 records. [Total pages: 111]". - Isolate the total pages segment:
substring-before(..., ']')captures everything up to the closing square bracket, resulting in"Showing 1 - 500 of 55385 records. [Total pages: 111". - Extract the number portion:
substring-after(..., 'Total pages: ')takes the text that follows the "Total pages: " label, giving us" 111"(there's a leading space here). - Clean up extra whitespace:
normalize-space()trims any leading or trailing spaces, leaving us with the clean numeric value"111".
Alternative (without normalize-space):
If you’re confident the spacing around the number will always be consistent, you can skip the normalize-space() step—just note this will return the number with a leading space:
substring-after(substring-before(//span[@id='ctl00_ContentPlaceHolder1_PagerControl1_litRecords']/text(), ']'), 'Total pages: ')
You’d then need to trim that space in your code if required, but using normalize-space() is more robust to any unexpected spacing changes in the future.
内容的提问来源于stack exchange,提问作者Rajendra Kunwar
相关产品推荐
相关产品推荐

