如何用Python抓取网页中特定文本后的标签内容?
问题描述
现有网页抓取代码通过固定位置的XPath定位span标签获取英文名,但网页结构可能变化,希望改为抓取文本**"english name:"**之后的<span>标签内容。
现有抓取代码:
# Find the span that contains the English name try: english_name_tag = driver.find_element(By.XPATH, '//*[@id="infoTag1"]/div/table/tbody/tr[1]/td[1]/span[2]') english_name = english_name_tag.text except: english_name = '-' # Find the rating tag try: rating_tag = driver.find_element(By.XPATH, '//*[@id="Div1"]/table/tbody/tr/td[1]/div[2]') rating_name = rating_tag.text except: rating_name = '-'
网页源码片段:
<div style="line-height: 24px; text-align: right" align="right"> <span id="QuickNav"> <a style="color:#3095b4;font-weight:bold;"><i class="fa fa-align-right" aria-hidden="true"></i>ניווט מהיר</a> <div id="CompanyNavBar" style="border:#3095b4 1px solid;border-bottom:0px; display: none; position: absolute; top: 145px;"> <div class="CompanyNav"><a href="#Div23">נתונים עסקיים</a></div> <div class="CompanyNav"><a href="#Div8">מנהלים</a></div> <div class="CompanyNav"><a href="#infoTag2">מודיעין עסקי</a></div> <div class="CompanyNav t5"><a href="#compDetailsDunsScore">דנסקור</a></div> <div class="CompanyNav"><a href="#compDetailsTreeStatic">עץ בעלויות</a></div> <div class="CompanyNav"><a href="#infoTag7">בעלות והנהלה</a></div> <div class="CompanyNav show_mobile"><a href="#neworder">כלים</a></div> </div> </span> <table width="100%" cellpadding="4" cellspacing="4"> <tbody> <tr> <td colspan="2"> <img src="/ShowImage?d_number=0" style="max-width: 210px; max-height: 90px; padding-bottom: 10px"> <div class="company_name"> עזריאלי אי קומרס בע"מ </div> known as: <span dir="ltr"> טיול במאדים </span> <br> second name: <span dir="ltr"> צטיילים במאדים </span> <br> english name: <span dir="ltr"> MARS TRIP </span> <br> <span class="comp_active">Active</span> <br> <br> </td> <td rowspan="1" valign="top" align="left" nowrap=""> <a style="cursor: pointer; text-decoration: none; display: inline; color: #aaaaaa; font-size: 13px"> <span style="cursor:pointer" onclick="ManageFollow(this)"> <img src="/img/star.png" class="favoritePng" style="width: 50px !important" duns="532450587" isinfollow="false" onclick="ManageFollow(this);updateAnalytics('event', 'Click', {'event_category': 'CompanysCard', 'event_label': 'AddMonitor'});"> <br> <span class="followTxt" duns="532450587">follow</span> </span> </a> </td> </tr> <tr> <td style="width: 20%; font-weight: bold; height: 30px">address</td> <td style="width: 69%; height: 30px">mars colony, 12345, mars</td> </tr> <tr> <td style="width: 20%; font-weight: bold; height: 30px">Phone</td> <td style="width: 69%; height: 30px;direction:ltr"> 123-456789 <br> </td> </tr> <tr> <td style="width: 20%; font-weight: bold; height: 30px">email</td> <td style="width: 69%; height: 30px"><a href="mailto:annonym@gmail.com">mailto:annonym@gmail.com</a></td> </tr> <tr> <td style="width: 20%; font-weight: bold; height: 30px"> Business category:</td> <td style="width: 69%; height: 30px">Food Industry, Packaging</td> </tr> </tbody> </table> </div>
解决方案
可以通过XPath的文本匹配+相邻节点选择实现需求,直接定位包含"english name:"的文本节点,再选择其后续的第一个<span>标签。
修改后的XPath表达式
//text()[normalize-space()="english name:"]/following-sibling::span[1]
normalize-space():去除文本前后的空格,避免因空格差异导致匹配失败following-sibling::span[1]:选择当前文本节点的第一个同级后续<span>标签
修改后的代码
# Find the span that contains the English name try: english_name_tag = driver.find_element(By.XPATH, '//text()[normalize-space()="english name:"]/following-sibling::span[1]') english_name = english_name_tag.text.strip() except: english_name = '-' # 原评分抓取代码保持不变 try: rating_tag = driver.find_element(By.XPATH, '//*[@id="Div1"]/table/tbody/tr/td[1]/div[2]') rating_name = rating_tag.text except: rating_name = '-'
兼容大小写的扩展写法
如果网页中"english name:"的文本存在大小写差异,可使用translate()实现不区分大小写的匹配:
//text()[normalize-space(translate(., 'ABCDEFGHIJKLMNOPQRSTUVWXYZ', 'abcdefghijklmnopqrstuvwxyz'))="english name:"]/following-sibling::span[1]
内容的提问来源于stack exchange,提问作者lidorkalfa
相关产品推荐
相关产品推荐

