使用Python Scrapy抓取adapt.io网站数据时所有字段返回空值如何解决
问题原因及修复方案
核心问题
- 你代码里使用的带随机后缀的类名是网站前端构建时自动生成的动态类名,每次部署都会变化,硬编码这类类名的选择器必然会匹配不到内容
- 跳转逻辑冗余错误:公司详情页本身就包含了所有公司基础信息、全部联系人信息,不需要额外跳转到联系人个人页面抓取,你跳转到联系人页后用的xpath完全不匹配该页面结构,自然返回空
- xpath语法错误:子节点查询没有以
.开头,导致每次都是全局搜索匹配,无法命中对应子元素 - 部分变量后多写了逗号,导致结果变成元组而不是字符串,即使匹配到也会出现格式错误
修正后可运行代码
import scrapy from urllib.parse import urlparse class CompanyProfileSpider(scrapy.Spider): name = 'companyDetails' start_urls = ["https://www.adapt.io/directory/industry/telecommunications/A-1"] custom_settings = { 'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } def parse(self, response): # 选公司列表项,不用动态类名,用结构定位 for company in response.xpath("//div[contains(@class, 'DirectoryList_linkItemWrapper')]/a"): name = company.xpath('./text()').get() company_portal = company.xpath('./@href').get() if company_portal: yield scrapy.Request( response.urljoin(company_portal), callback=self.company_parse, meta={'company_name': name} ) def company_parse(self, response): # 提取公司基础信息,不用动态类名,用标签内容定位 base_info = {} for item in response.xpath("//div[contains(@class, 'CompanyTopInfo_contentWrapper')]"): key = item.xpath('./span[1]/text()').get('').strip() value = item.xpath('./span[2]/text()').get('').strip() base_info[key] = value website = base_info.get('Website', '') # 提取域名 domain = urlparse(website).netloc if website else '' # 提取所有联系人 contact_list = [] for contact in response.xpath("//div[contains(@class, 'TopContacts_roundedBorder')]"): contact_name = contact.xpath(".//div[contains(@class, 'TopContacts_contactName')]//a/text()").get('').strip() jobtitle = contact.xpath(".//p[contains(@class, 'TopContacts_jobTitle')]/text()").get('').strip() department = contact.xpath(".//p[contains(@class, 'TopContacts_department')]/text()").get('').strip() contact_list.append({ "contact_name": contact_name, "contact_jobtitle": jobtitle, "contact_email_domain": domain, "contact_department": department if department else "Other" }) # 按要求的格式输出 yield { "company name": response.meta.get('company_name', ''), "company_location": base_info.get('Location', None), "company_website": website if website else None, "company_webdomain": domain if domain else None, "company_industry": base_info.get('Industry', None), "company_employee_size": base_info.get('Head Count', None), "company_revenue": base_info.get('Revenue', None), "contact_details": contact_list }
使用说明
- 代码新增了
USER_AGENT配置,避免被网站反爬拦截返回空页面 - 完全去掉了动态类名的随机后缀匹配,只匹配类名的固定前缀,稳定性更高
- 直接在公司详情页提取所有需要的信息,删除了冗余的联系人页跳转逻辑
- 自动根据公司网址提取域名填充到联系人的
contact_email_domain字段,符合要求的输出格式 - 空值自动设为
None,完全匹配你期望的输出结构
内容的提问来源于stack exchange,提问作者grroom train
相关产品推荐
相关产品推荐

