使用BeautifulSoup时a:nth-of-type(n)语法的疑问求助
关于BeautifulSoup中
p ~ a:nth-of-type(n)选择器的疑问 作为BeautifulSoup新手,我遇到了一个选择器的问题,相关代码如下:
html_doc = """ <html><head><title>"happy new year"</title></head> <body> <p class="title"><b>c.biancheng.net</b></p> <p class="website">2022 is last year</p> <a href="http://c.biancheng.net/python/" id="link1">2023 is the new year</a> <a href="http://c.biancheng.net/c/" id="link2">hope you will be better</a> <a href="http://c.biancheng.net/django/" id="link3">achieve more</a> <p class="VIP">YEAH</p> <a href="http://vip.biancheng.net/?from=index" id="link4">and realize your dream</a> <p class="introduce">introduction: <a href="http://c.biancheng.net/view/8066.html" id="link5">best</a> <a href="http://c.biancheng.net/view/8092.html" id="link6">wishes</a> </p> """ soup = BeautifulSoup(html_doc, 'html.parser') print(soup.select('p ~ a:nth-of-type(4)'))
执行上述代码会输出:
[<a href="http://vip.biancheng.net/?from=index" id="link4">and realize your dream</a>]
但将代码中的4改为5,即执行print(soup.select('p ~ a:nth-of-type(5)'))时,返回值为[](空列表),我不清楚原因,希望有人能解答这一现象。
问题解答
要理解这个现象,得拆解选择器的逻辑:
p ~ a:只匹配在<p>标签之后、且与<p>同级的<a>元素。你的HTML里,link1、link2、link3、link4都是<body>的直接子元素,和前面的<p>同级,会被选中;而link5、link6是<p class="introduce">的子元素,属于嵌套关系,不是<p>的同级兄弟,所以不会被这个选择器匹配到。:nth-of-type(n):这个伪类是以元素的父元素为范围来计数同类型元素的。这里匹配到的<a>都属于<body>的子元素,而<body>下总共只有4个符合条件的<a>,所以当n=5时,找不到对应的元素,返回空列表。
当n=4时,正好对应<body>下的第4个<a>(也就是link4),所以能正常输出结果。
内容的提问来源于stack exchange,提问作者TryingToLearnSomething
相关产品推荐
相关产品推荐

