Scrapy爬取指定表格失败求助:无法获取equityCompleteHoldingTable数据
解决Scrapy提取目标表格数据的问题
问题分析
你当前的XPath错误在于,使用table.xpath('//tr')时,//会从整个文档根节点重新查找所有<tr>元素,而非在目标表格equityCompleteHoldingTable的范围内查找,导致提取到了其他表格的行数据。
修正后的Scrapy操作代码
# 启动Scrapy Shell scrapy shell 'https://www.moneycontrol.com/mutual-funds/canara-robeco-blue-chip-equity-fund-direct-plan/portfolio-holdings/MCA212' # 定位目标表格(注意ID的XPath写法,直接用单引号包裹ID值即可) table = response.xpath('//*[@id="equityCompleteHoldingTable"]') # 在目标表格范围内筛选有效数据行(.//表示从当前节点向下查找) rows = table.xpath('.//tr[contains(@class, "row")]') # 提取单条数据示例 row = rows[0] stock_name = row.xpath('.//td[1]//text()').get().strip() sector = row.xpath('.//td[2]//text()').get().strip() # 批量转换为字典格式 stock_data_list = [] for row in rows: stock_dict = { '股票名称': row.xpath('.//td[1]//text()').get(default='').strip(), '所属行业': row.xpath('.//td[2]//text()').get(default='').strip(), '持仓占比': row.xpath('.//td[3]//text()').get(default='').strip(), '市值(亿)': row.xpath('.//td[4]//text()').get(default='').strip(), '持有数量': row.xpath('.//td[5]//text()').get(default='').strip(), '平均成本': row.xpath('.//td[6]//text()').get(default='').strip() } stock_data_list.append(stock_dict) # 查看转换后的字典数据 print(stock_data_list)
关键修正点
- 使用
.//替代//:.//表示从当前节点(目标表格)开始向下查找,避免全局匹配其他表格的元素 - 修正ID的XPath写法:直接用
[@id="equityCompleteHoldingTable"],无需转义引号 - 过滤有效行:通过
contains(@class, "row")筛选表格的数据行,排除表头和汇总行
内容的提问来源于stack exchange,提问作者Learning Product Management
相关产品推荐
相关产品推荐

