Zefix搜索已有条目无结果问题排查及法律形式提取
Zefix平台搜索结果无法获取的代码错误排查
我希望从Zefix平台提取搜索词“Fondazione Amici di AMCA”的法律形式,预期输出为对应详情页中的“Foundation”,但运行以下Python代码时始终返回“No search results found.”,尽管该条目确实存在,请排查代码中的错误:
import requests from bs4 import BeautifulSoup search_url = "https://www.zefix.ch/de/search/entity/welcome" search_params = {"name": "Fondazione Amici di AMCA"} response = requests.get(search_url, params=search_params) soup = BeautifulSoup(response.content, "html.parser") # Find the link to the details page for the first search result search_results = soup.find_all("div", {"class": "result-entry"}) if search_results: first_result = search_results[0] details_link = first_result.find("a", {"class": "list-group-item-1"}) if details_link: details_url = f"https://www.zefix.ch{details_link['href']}" # Visit the details page and extract the legal form details_response = requests.get(details_url) details_soup = BeautifulSoup(details_response.content, "html.parser") legal_form = details_soup.find("td", text="Rechtsform:").find_next_sibling("td").get_text(strip=True) print(f"The legal form of Fondazione Amici di AMCA is: {legal_form}") else: print("No link to details page found.") else: error_message = soup.find("div", {"class": "no-result-message"}) if error_message: print(error_message.text.strip()) else: print("No search results found.")
代码错误分析
- 语言版本不匹配:代码中使用了德语版搜索URL (
/de/),但目标条目对应的是英语版页面。Zefix的多语言页面是独立索引的,德语版搜索可能无法匹配到该条目,或者返回结果的DOM结构与英语版不一致。 - 详情页文本匹配错误:代码中用德语关键词
Rechtsform:查找法律形式,但目标详情页是英语版,对应关键词应为Legal form:,即便能进入详情页也会提取失败。
修正后的代码
import requests from bs4 import BeautifulSoup # 使用英语版搜索页面,与目标详情页语言保持一致 search_url = "https://www.zefix.ch/en/search/entity/welcome" search_params = {"name": "Fondazione Amici di AMCA"} response = requests.get(search_url, params=search_params) soup = BeautifulSoup(response.content, "html.parser") # 查找搜索结果条目 search_results = soup.find_all("div", {"class": "result-entry"}) if search_results: first_result = search_results[0] details_link = first_result.find("a", {"class": "list-group-item-1"}) if details_link: details_url = f"https://www.zefix.ch{details_link['href']}" # 访问详情页并提取法律形式(使用英语关键词匹配) details_response = requests.get(details_url) details_soup = BeautifulSoup(details_response.content, "html.parser") legal_form = details_soup.find("td", text="Legal form:").find_next_sibling("td").get_text(strip=True) print(f"Fondazione Amici di AMCA的法律形式是: {legal_form}") else: print("未找到详情页链接。") else: error_message = soup.find("div", {"class": "no-result-message"}) if error_message: print(error_message.text.strip()) else: print("未找到搜索结果。")
内容的提问来源于stack exchange,提问作者David
相关产品推荐
相关产品推荐

