Python BeautifulSoup爬取Bruttofortjeneste字段失败求助
解决Bruttofortjeneste字段爬取问题
你当前代码无法获取Bruttofortjeneste值,是因为该字段的DOM结构与另外两个字段不同——它位于页面顶部的统计卡片区域,而非下方的详情标签组,原查找逻辑的标签匹配规则不适用。
修正后的代码如下:
import requests from bs4 import BeautifulSoup url = 'https://ownr.dk/companies/public-profile/34883793' response = requests.get(url) soup = BeautifulSoup(response.content, 'html.parser') # 修正Bruttofortjeneste的查找逻辑 bruttofortjeneste_value = 'N/A' # 通过文本模糊匹配找到目标标签,避免空格或标签结构差异影响 brutto_label = soup.find('div', string=lambda text: text and 'Bruttofortjeneste' in text.strip()) if brutto_label: # 获取相邻的数值元素 brutto_value_elem = brutto_label.find_next_sibling('div') if brutto_value_elem: bruttofortjeneste_value = brutto_value_elem.text.strip() # 保留原有的Antal ansatte和Branchekode获取逻辑 antal_ansatte_elem = soup.find('div', {'class': 'label'}, string='Antal ansatte') antal_ansatte_value = antal_ansatte_elem.find_next_sibling('div').text.strip() if antal_ansatte_elem else 'N/A' branchekode_elem = soup.find('div', {'class': 'label'}, string='Branchekode') branchekode_value = branchekode_elem.find_next_sibling('div').text.strip() if branchekode_elem else 'N/A' print('Bruttofortjeneste:', bruttofortjeneste_value) print('Antal ansatte:', antal_ansatte_value) print('Branchekode:', branchekode_value)
说明
改用lambda表达式模糊匹配包含目标文本的元素,能避免因标签class、文本前后空格导致的匹配失败;找到标签后再获取其相邻的数值元素,即可正确拿到Bruttofortjeneste的值。
内容的提问来源于stack exchange,提问作者Pythonnoob
相关产品推荐
相关产品推荐

