如何使用Beautiful Soup遍历URL列表爬取多个文本字段
问题修复方案
现有错误原因
- 语法错误:遍历
psoup2.find('p','col-lg-7 lh14e').text的代码行末尾缺少英文冒号,触发语法报错 - 逻辑错误:
psoup2.find('p','col-lg-7 lh14e').text已经直接拿到了姓名字符串,你额外遍历字符串的每个字符,还调用不存在的get('text')方法,触发属性不存在报错。
嵌套循环必要性说明
不需要嵌套循环。后续爬取年龄、俱乐部等其他字段时,只需要在遍历单页的同一层级循环内,对应HTML结构写定位规则即可。
修复后代码示例
基础修复版(仅爬取姓名)
names = [] # 遍历球员链接列表 for link in playerLinks: reqs2 = Request(link, headers=headers) html_page = urlopen(reqs2) psoup2 = BeautifulSoup(html_page, "html.parser") # 直接复用单页测试的正确逻辑,加判空避免部分页面结构异常报错 name_elem = psoup2.find('p','col-lg-7 lh14e') if name_elem: names.append(name_elem.text)
优化版(支持多字段爬取)
对应你给出的HTML结构,所有字段都在row cf lightGrayBorderBottom容器内,左侧为字段名,右侧为字段值,可以写通用逻辑批量取值:
player_data = [] for link in playerLinks: reqs2 = Request(link, headers=headers) html_page = urlopen(reqs2) psoup2 = BeautifulSoup(html_page, "html.parser") # 存储单名球员的所有字段 per_player = {} # 遍历页面上所有信息行 for row in psoup2.find_all('div', class_='row cf lightGrayBorderBottom'): field = row.find('p', class_='col-lg-5').text.strip() value = row.find('p', class_='col-lg-7').text.strip() per_player[field] = value player_data.append(per_player) # 后续可直接从结果中读取需要的字段 print(player_data[0].get('Full Name')) print(player_data[0].get('Age')) print(player_data[0].get('Club'))
内容的提问来源于stack exchange,提问作者SlowBear
相关产品推荐
相关产品推荐

