如何用Python的BeautifulSoup提取表格列中的href链接
代码修改方案
你需要在拿到每行的td列表后,定位到包含目标a标签的td,再提取href属性即可,修改后的完整代码如下:
from bs4 import BeautifulSoup import requests try: r = requests.get('https://coinmarketcap.com/') soup = BeautifulSoup(r.text, 'lxml') table = soup.find('table', class_='cmc-table') for row in table.tbody.find_all('tr'): columns = row.find_all('td') # 先判断列数足够,避免索引报错 if len(columns) >=3: # 定位第三个td里的cmc-link类a标签 target_a = columns[2].select_one('a.cmc-link') if target_a: # 提取href属性 href = target_a.get('href') print(href) except requests.exceptions.RequestException as e: print(e)
关键修改说明
- 你给出的示例列数据中,包含加密货币链接的是每行的第三个
<td>元素,对应列表索引为2,直接取该列做后续筛选 - 使用
select_one('a.cmc-link')可以快速匹配指定类名的a标签,避免逐层查找元素的冗余代码 - 使用
get('href')获取属性值,当匹配不到对应元素时会返回None,不会直接抛出异常中断程序运行
内容的提问来源于stack exchange,提问作者Ehsan Akbar
相关产品推荐
相关产品推荐

