You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取网页表格文本与URL遇阻:仅能提取文本求方案

解决BeautifulSoup提取表格中文本与URL的问题

嘿,我懂你现在的烦恼——用BeautifulSoup爬网站表格时,只能抽取出纯文本,想要同时保留链接的URL却搞不定,而且你也猜到问题出在text.strip()上对吧?确实,text方法只会提取标签里的纯文本内容,把<a>标签的href属性直接丢了。咱们来改改代码,让它同时拿到文本和对应的URL。

核心思路

表格里的链接都在<a>标签里,所以咱们不能直接拿单元格的文本,得先检查单元格里有没有<a>标签:

  • 如果有,就分别提取标签的文本(link.text.strip())和href属性(link.get('href'))
  • 如果没有,就只提取单元格的纯文本,同时把URL设为None保持数据结构一致

修改后的完整代码

import requests
from bs4 import BeautifulSoup
from requests.compat import urljoin  # 用来处理相对路径转绝对URL

start_number = 0
max_number = 5
results = []  # 用这个列表存每行的文本+URL组合,比单独存urls更实用

for number in range(start_number, max_number + start_number):
    # 补全你的目标URL,这里假设是带number参数的地址
    base_url = 'http://www.ispo-org.or.id/index.php?op=...'
    target_url = f'{base_url}&num={number}'
    
    try:
        response = requests.get(target_url)
        response.raise_for_status()  # 检查请求是否成功
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # 精准定位目标表格,比如根据id或class,避免拿到无关表格
        target_table = soup.find('table', class_='your-table-class')  # 替换成你表格的class/id
        if not target_table:
            print(f"页面{target_url}中未找到目标表格,跳过")
            continue
        
        # 遍历表格的每一行
        for row in target_table.find_all('tr'):
            row_content = []
            # 遍历每行的单元格(td或th)
            for cell in row.find_all(['td', 'th']):
                link = cell.find('a')
                if link:
                    # 提取链接文本和URL,相对路径转绝对URL
                    link_text = link.text.strip()
                    link_url = urljoin(base_url, link.get('href'))  # 处理相对路径
                    row_content.append({'显示文本': link_text, '链接地址': link_url})
                else:
                    # 无链接时只存文本
                    row_content.append({'显示文本': cell.text.strip(), '链接地址': None})
            results.append(row_content)
    
    except requests.exceptions.RequestException as e:
        print(f"请求页面{target_url}失败: {str(e)}")

# 打印结果示例
for i, row in enumerate(results, 1):
    print(f"\n第{i}行数据:")
    for item in row:
        print(f"文本: {item['显示文本']} | URL: {item['链接地址']}")

关键细节说明

  • 处理相对路径:如果网站里的链接是相对路径(比如/page.php),用urljoin(base_url, href)可以把它转换成完整的绝对URL,直接就能访问。
  • 精准定位表格:一定要用find('table', class_='xxx')或find('table', id='xxx')定位目标表格,不然可能会爬取到页面里其他无关的表格数据。
  • 异常处理:加了try-except块处理请求异常,避免某个页面请求失败导致整个程序崩溃。

这样改完之后,你就能同时拿到表格里的文本和对应的URL啦!

内容的提问来源于stack exchange,提问作者Funkeh-Monkeh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:31:18