使用Kanna解析HTML时获取图片URL遇阻,求解决方案
data-high-quality from the Second img Element with Kanna Hey there! Let's work through how to grab that second img element's data-high-quality attribute you need. Here are a couple of reliable approaches you can use:
1. Directly Target the Second img with CSS Selector
Use the :nth-of-type(2) pseudo-class to pick the second img element in its parent container. This is straightforward if you know the structure of your HTML:
import Kanna // Replace yourHTMLString with your actual HTML content if let doc = try? HTML(html: yourHTMLString, encoding: .utf8) { // Select the second img element if let secondImg = doc.css("img:nth-of-type(2)").first { // Extract the data-high-quality attribute if let highQualityURL = secondImg["data-high-quality"] { print("Found high-quality URL: \(highQualityURL)") } else { print("The second img doesn't have a data-high-quality attribute.") } } else { print("Couldn't find a second img element in the document.") } }
2. Use XPath to Select the Second img
XPath offers another way to target the second img element directly with the position() function:
import Kanna if let doc = try? HTML(html: yourHTMLString, encoding: .utf8) { // Target the second img element using XPath if let secondImg = doc.xpath("//img[position()=2]").first { if let highQualityURL = secondImg["data-high-quality"] { print("Found high-quality URL: \(highQualityURL)") } } }
3. Iterate Through All img Elements (Safer for Dynamic Content)
If your HTML structure might vary, it's safer to fetch all img elements first, check if there are at least two, then access the second one (remember Swift arrays are 0-indexed, so we use index 1):
import Kanna if let doc = try? HTML(html: yourHTMLString, encoding: .utf8) { let allImages = doc.css("img") // Make sure there are at least 2 img elements guard allImages.count >= 2 else { print("Not enough img elements found.") return } let secondImage = allImages[1] if let highQualityURL = secondImage["data-high-quality"] { print("Found high-quality URL: \(highQualityURL)") } else { print("Second img element has no data-high-quality attribute.") } }
Each method works depending on your HTML structure—if the second img is always the second child of its parent, the first two methods are quick. If you need to handle cases where the number of images might vary, the third method is more robust.
内容的提问来源于stack exchange,提问作者matteo94

