You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup+Python表格遍历重复读取及链接修复问题求助

问题1:循环重复读取同一行数据

现象

运行代码时持续输出同一行囚犯的信息,示例如下:

https://www.tdcj.texas.gov/death_row/dr_info/murphyjeddidiahlast.html
--- retrieving statement 584  ---
--- retrieving execution data for execution ID 584 Murphy , Jeddidiah  ---
...(重复内容)

原因

你的循环逻辑完全错误:

  • 提前用data = table.find_all('td')取出了所有单元格,循环里却一直用data[0]、data[3]这类固定索引,每次都只取第一行的内容。
  • 遍历rows时,没有针对当前行提取单元格,而是依赖全局的data变量。
  • 获取最后陈述链接时,LastStatementLinks = table.find_all("a", href=True)取的是整个表格的所有链接,再固定取LastStatementLinks[1],自然每次都是同一个链接。

修复代码

遍历行时,只提取当前行的单元格,且从当前行内获取链接:

# 移除全局的data = table.find_all('td')
rows = table.find_all('tr')
for row in rows[1:]:  # 跳过表头行
    cells = row.find_all('td')  # 仅获取当前行的单元格
    ExecutionID = str(cells[0].get_text())
    Lastname = str(cells[3].get_text())
    Firstname = str(cells[4].get_text())
    TDJC = str(cells[5].get_text())
    Age = str(cells[6].get_text())
    Date = str(cells[7].get_text())
    Race = str(cells[8].get_text())
    County = str(cells[9].get_text())
    
    # 从当前行获取最后陈述链接
    last_statement_a = cells[2].find('a', href=True)
    if last_statement_a:
        LastStatementLink = last_statement_a.get("href")
        Urlcomplete = UrlLastStatement + LastStatementLink
        print(Urlcomplete)
    
    # ...后续获取陈述、插入数据库的代码

问题2:部分囚犯最后陈述链接错误

现象

部分囚犯(如545、544、552)的链接被错误拼接,例如囚犯545的链接变成https://www.tdcj.texas.gov/death_row/death_row/dr_info/cardenasrubenlast.html,导致无法访问。

原因

  • 赋值错误:if语句里用了==(比较运算符)而非=(赋值运算符),比如linkLS == "xxx"根本没修改变量值。
  • 硬编码修复逻辑有缺陷:修改特殊ID的链接后,没有执行后续的陈述获取逻辑,导致这几个ID的陈述内容为空。

修复方案

方案1:统一处理链接(推荐)

使用urljoin自动处理路径重复问题,无需针对单个ID硬编码:

from urllib.parse import urljoin

base_url = 'https://www.tdcj.texas.gov'
# ...
for row in table.find_all("tr")[1:]:
    cells = row.find_all('td')
    # ...其他字段提取
    # 用urljoin拼接链接,自动处理重复前缀
    linkLS = urljoin(base_url, cells[2].a['href'])

方案2:修正硬编码赋值错误

如果坚持用ID判断修复,要修正赋值符号,并且统一执行陈述获取逻辑:

for row in table.find_all("tr")[1:]:
    cells = row.find_all('td')
    ExecutionID= str(cells[0].get_text())
    # ...其他字段提取
    linkLS = 'https://www.tdcj.texas.gov/death_row/' + cells[2].a['href']
    
    # 修正特殊ID的链接
    if ExecutionID == "545":
        linkLS = "https://www.tdcj.texas.gov/death_row/dr_info/cardenasrubenlast.html"
    elif ExecutionID == "544":
        linkLS = "https://www.tdcj.texas.gov/death_row/dr_info/pruettrobertlast.html"
    elif ExecutionID == "552":
        linkLS = "https://www.tdcj.texas.gov/death_row/dr_info/no_last_statement.html"
    
    # 统一执行陈述获取逻辑
    try:
        response = requests.get(linkLS, verify=False)
        print ("--- retrieving statement", ExecutionID," ---")
        print (linkLS)
        soup = BeautifulSoup(response.text, "html.parser")
        pattern = re.compile("Last Statement:")
        statement = soup.find(string=pattern).findNext('p').contents[0]
        # ...清理陈述内容的代码
    except Exception:
        print("An error has occured while retrieving the statement. Proceeding...")
        statement = ""
    
    # ...插入数据库的代码

额外优化

  1. 移除冗余的urllib.request.urlopen,直接用requests.get返回的响应文本解析。
  2. 修正表格创建语句的语法错误:InmateInformation, text改为InmateInformation text(缺少空格)。

内容的提问来源于stack exchange,提问作者Omnicron

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 21:54:56