You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests模块无法抓取网站GRANTOR与GRANTEE列名称问题

问题排查与解决

可能的原因及对应方案

1. 页面为动态渲染,requests无法获取异步加载的内容

该网站的搜索结果大概率是通过JavaScript动态加载的,requests只能抓取初始静态HTML,拿不到后续渲染生成的表格数据。

解决方法:使用selenium或playwright模拟浏览器完整加载页面,示例代码(selenium):

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

link = 'https://dallas.tx.publicsearch.us/results?_docTypes=AMD%2CDT%2CMOD%2CMODT%2CDE%2CCORR%20DT&department=RP&limit=50&offset=0&recordedDateRange=20181101%2C20231010&searchOcrText=true&searchType=quickSearch&searchValue=%22adjustable%20rate%20rider%22'

# 初始化浏览器(需提前安装对应浏览器驱动)
driver = webdriver.Chrome()
driver.get(link)

# 等待表格加载完成
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "[data-tourid='searchResults'] table")))

# 获取完整页面源码解析
soup = BeautifulSoup(driver.page_source, "lxml")
driver.quit()

for item in soup.select("[data-tourid='searchResults'] table tr"):
    tds = item.select("td")
    if len(tds) >= 5:  # 确保存在目标列
        grantor = tds[3].text.strip()
        grantee = tds[4].text.strip()
        print(grantor, grantee)

2. CSS选择器或列索引错误

  • 很多网站的表格tbody标签是浏览器渲染时自动添加的,原始HTML中并不存在,导致table > tbody > tr选择不到元素。可以去掉tbody,直接用table tr作为选择器。
  • 列索引可能有误,建议先打印所有单元格文本,确认GRANTOR和GRANTEE的实际位置:
for item in soup.select("[data-tourid='searchResults'] table tr"):
    tds = item.select("td")
    print([td.text.strip() for td in tds])

3. 请求头不足被网站拦截

尝试补充更多请求头字段,模拟真实浏览器的请求特征:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/117.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://dallas.tx.publicsearch.us/',
    'Connection': 'keep-alive',
}

4. 先验证请求是否成功

在代码中添加基础检查,确认是否获取到正确的页面内容:

res = requests.get(link, headers=headers)
print(res.status_code)  # 正常应返回200
print(res.text[:1000])  # 查看返回内容的前1000字符,确认是否包含搜索结果表格

内容的提问来源于stack exchange,提问作者robots.txt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 20:22:10