如何使用BeautifulSoup提取<p>标签中的末尾租金字符串?
解决BeautifulSoup爬取p标签内租金的问题
你的问题出在目标<p>标签包含了<strong>子标签和独立的文本节点:直接用price.string会返回None(因为标签有多个子节点时,string属性仅在标签只有一个文本子节点时生效);第二段代码只获取了<strong>的内容,没触达后面的租金文本。下面给你几种可行的解决方案:
方法1:用get_text()提取并处理文本
先获取<p>标签的全部文本,再通过字符串操作分离出租金:
text = soup.find_all("div", {"class": "plan-group rent"}) for item in text: rent_p = item.find("p") # 用find代替find_all,单个div下通常只有一个目标p标签 if rent_p: full_text = rent_p.get_text(strip=True) # 分割文本提取租金部分 rent = full_text.split("Monthly Rent:")[-1].strip() print(rent)
方法2:定位<strong>后取其兄弟节点
直接找到<strong>标签,获取它的下一个兄弟文本节点(就是租金内容):
text = soup.find_all("div", {"class": "plan-group rent"}) for item in text: strong_tag = item.find("strong", class_="hidden show-mobile-inline") if strong_tag: # next_sibling指向strong标签后面的文本内容 rent = strong_tag.next_sibling.strip() print(rent)
方法3:遍历<p>的子节点筛选文本
遍历<p>标签的所有子节点,提取非标签类型的纯文本内容:
text = soup.find_all("div", {"class": "plan-group rent"}) for item in text: rent_p = item.find("p") if rent_p: # 筛选出不是标签的子节点(即纯文本节点) for child in rent_p.children: if not child.name: rent = child.strip() if rent: # 排除空文本 print(rent)
另外注意你代码里的小错误:find_all("div", {"class", "plan-group rent"})中的字典写法有误,应该是{"class": "plan-group rent"}(用冒号而非逗号),否则无法正确匹配class属性。
内容的提问来源于stack exchange,提问作者Tyler Benton
相关产品推荐
相关产品推荐

