使用Beautiful Soup定位无类名<p>标签实现网页抓取
定位带类名
<p>后的无类名<p>标签的方法 这里有两种简单直接的方法可以实现需求:
方法一:使用find_next_sibling()方法
先找到带类名的<p>标签,再调用find_next_sibling()方法获取它后面紧邻的<p>标签:
from bs4 import BeautifulSoup html = '''<li class="pp-property-box"> <h2>Address</h2> <p class="pp-property-price">£$150,000></p> <p>5 bedroom mansion</p></li>''' soup = BeautifulSoup(html, 'html.parser') for li in soup.find_all('li', class_="pp-property-box"): title = li.find('h2').text price = li.find('p', class_="pp-property-price").text # 获取带类名p之后的无类名p标签文本 property_desc = li.find('p', class_="pp-property-price").find_next_sibling('p').text print(title, price, property_desc)
方法二:使用CSS相邻兄弟选择器+
利用CSS选择器的+语法,直接选中.pp-property-price类的<p>标签后面紧邻的<p>标签:
from bs4 import BeautifulSoup html = '''<li class="pp-property-box"> <h2>Address</h2> <p class="pp-property-price">£$150,000></p> <p>5 bedroom mansion</p></li>''' soup = BeautifulSoup(html, 'html.parser') for li in soup.find_all('li', class_="pp-property-box"): title = li.find('h2').text price = li.find('p', class_="pp-property-price").text # 用CSS选择器直接定位目标标签 property_desc = li.select_one('.pp-property-price + p').text print(title, price, property_desc)
两种方法都能精准获取到目标无类名<p>标签的内容,根据自己的习惯选择即可。
内容的提问来源于stack exchange,提问作者cts
相关产品推荐
相关产品推荐

