如何用Selenium和Beautiful Soup提取含aria-hidden的span文本?
解决方法:提取目标文本的BeautifulSoup参数填写
根据你提供的HTML结构,有两种直接的参数填写方式可以定位到目标元素:
方式一:定位外层span元素
外层span的class属性是t-14 t-normal,直接通过这个class定位:
text_loc = intro.find('span', {'class': 't-14 t-normal'})
方式二:定位内层带aria-hidden属性的span
内层span带有aria-hidden="true"属性,这个元素的文本就是你需要的内容,定位更精准:
text_loc = intro.find('span', {'aria-hidden': 'true'})
完整处理代码
定位后提取文本时,需要把原文本中的·替换为空格,得到你要的Crédit Agricole CIB Full-time:
src = driver.page_source soup = BeautifulSoup(src, 'lxml') intro = soup.find('div', {'class': 'pv-text-details__left-panel'}) # 选其中一种定位方式即可 text_loc = intro.find('span', {'aria-hidden': 'true'}) # 或者 text_loc = intro.find('span', {'class': 't-14 t-normal'}) text = text_loc.get_text().strip().replace(' · ', ' ')
内容的提问来源于stack exchange,提问作者LaC
相关产品推荐
相关产品推荐

