如何使用XPath提取特定数值?——从文本中提取数字的实例咨询
Got it, let's figure out how to pull just the number "6" from the text "ตอนที่ 6" you've already fetched with //li[@class='active']/text(). The approach depends on which XPath version your tool/library supports—here are the most common solutions:
XPath 2.0+ (Modern Tools/Libraries)
If you're using an XPath 2.0 or newer compatible tool (like Saxon, or modern scrapers that support it), you have flexible regex-based options:
Strip all non-digit characters with
replace():replace(//li[@class='active']/text(), '[^\d]', '')This regex replaces every character that isn't a digit (
[^\d]) with an empty string, leaving only the number behind.Split the string and grab the last segment with
tokenize():
Since the number is at the end of the string separated by a space, you can split on whitespace and pick the last item:tokenize(//li[@class='active']/text(), '\s+')[last()]\s+matches one or more whitespace characters, so even if there are multiple spaces, this will still work.
XPath 1.0 (Older Tools/Libraries)
XPath 1.0 doesn't support regex or string splitting, but we can use basic string functions since your target text has a consistent format:
Use
substring-after()to get text after the space:substring-after(//li[@class='active']/text(), ' ')This works because the number comes right after a single space in "ตอนที่ 6".
Handle potential extra spaces with
normalize-space():
If there's a chance of multiple spaces between "ตอนที่" and the number, normalize the whitespace first:substring-after(normalize-space(//li[@class='active']/text()), ' ')normalize-space()collapses multiple spaces into one, trims leading/trailing spaces, so thesubstring-afterwill reliably find the number.
All these methods will give you either "6" (or " 6 " if there are trailing spaces, but normalize-space() can fix that too if needed!).
内容的提问来源于stack exchange,提问作者parames3010

