自定义优先级输出失败及网页链接提取优先级异常问题求助
问题1:按指定优先级输出列表元素
你的代码只是遍历原列表逐个打印,完全没改变输出顺序。要实现E、D、B、A的固定优先级输出,得先定义好目标顺序,再按这个顺序去原列表里匹配元素打印。
修改后的代码:
items = ['A','B','D','E'] # 定义优先级顺序 priority_order = ['E', 'D', 'B', 'A'] for item in priority_order: if item in items: print(item)
不管原列表元素顺序如何,都会严格按照priority_order的顺序输出存在的元素。
问题2:网页链接提取优先级失效
你的遍历逻辑有bug:遇到第一个不满足contact的元素时,只要它满足about就会直接赋值并跳出循环,根本没机会往后查找contact链接。比如页面里第一个a标签是about,程序就直接返回了,哪怕后面有contact链接。
正确逻辑是:先找contact链接,找到就终止查找;完全找不到contact时,再去查找about链接。
方案1:用CSS选择器直接定位(更简洁)
import requests from urllib.parse import urljoin from bs4 import BeautifulSoup links = [ "http://www.innovaprint.com.sg/", "https://www.richardsonproperties.com/", "https://www.thepunctuationguide.com/", "http://www.knowledgeplatform.com/", "http://www.singaporeenterpriseassociation.com/", "https://www2.deloitte.com/sg/en.html" ] headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:105.0) Gecko/20100101 Firefox/105.0' } for link in links: res = requests.get(link, headers=headers) soup = BeautifulSoup(res.text, "lxml") target_link = '' # 优先查找contact链接 contact_item = soup.select_one('a[href]:contains("contact")') if contact_item: target_link = urljoin(link, contact_item['href']) else: # 没找到contact再找about about_item = soup.select_one('a[href]:contains("about")') if about_item: target_link = urljoin(link, about_item['href']) print(target_link)
方案2:分两次遍历(兼容性更强)
如果:contains选择器不生效,就用两次遍历实现优先级:
import requests from urllib.parse import urljoin from bs4 import BeautifulSoup links = [ "http://www.innovaprint.com.sg/", "https://www.richardsonproperties.com/", "https://www.thepunctuationguide.com/", "http://www.knowledgeplatform.com/", "http://www.singaporeenterpriseassociation.com/", "https://www2.deloitte.com/sg/en.html" ] headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:105.0) Gecko/20100101 Firefox/105.0' } for link in links: res = requests.get(link, headers=headers) soup = BeautifulSoup(res.text, "lxml") target_link = '' # 第一次遍历:找contact for item in soup.select("a[href]"): if "contact" in item.text.lower(): target_link = urljoin(link, item['href']) break # 没找到contact,第二次遍历找about if not target_link: for item in soup.select("a[href]"): if "about" in item.text.lower(): target_link = urljoin(link, item['href']) break print(target_link)
内容的提问来源于stack exchange,提问作者SMTH
相关产品推荐
相关产品推荐

