You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于BeautifulSoup结果创建DataFrame失败,求提取链接文本方案

解决方法

错误原因

pd.read_html() 仅能解析HTML中的<table>标签结构,你当前获取的是<h3>标签集合,不存在表格元素,因此触发ValueError: No tables found。

正确实现步骤

无需使用pd.read_html(),直接从<h3>下的<a>标签提取链接和文本,再构造DataFrame即可:

from bs4 import BeautifulSoup
import requests
import pandas as pd
import lxml

def getHTMLdocument(url):
    response = requests.get(url) 
    return response.text

url_to_scrape = "https://website.com"
html_document = getHTMLdocument(url_to_scrape)
soup = BeautifulSoup(html_document, 'lxml')
h3_list = soup.find_all('h3')

# 提取链接与文本数据
data = []
for h3 in h3_list:
    a_tag = h3.find('a', class_='xyz-link')
    if a_tag:  # 避免找不到a标签引发报错
        link = a_tag.get('href')
        text = a_tag.get_text(strip=True)
        data.append({'链接地址': link, '显示文本': text})

# 生成DataFrame
df = pd.DataFrame(data)
print(df)

关键说明

  • 遍历<h3>标签集合,通过find()定位内部的目标<a>标签
  • 用get('href')获取链接属性值,get_text(strip=True)提取并去除文本首尾空白
  • 将每组数据存入列表,最后通过pd.DataFrame()直接构造结构化数据框

内容的提问来源于stack exchange,提问作者question12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 10:42:47