BeautifulSoup+Python表格遍历重复读取及链接修复问题求助
问题1:循环重复读取同一行数据
现象
运行代码时持续输出同一行囚犯的信息,示例如下:
https://www.tdcj.texas.gov/death_row/dr_info/murphyjeddidiahlast.html --- retrieving statement 584 --- --- retrieving execution data for execution ID 584 Murphy , Jeddidiah --- ...(重复内容)
原因
你的循环逻辑完全错误:
- 提前用
data = table.find_all('td')取出了所有单元格,循环里却一直用data[0]、data[3]这类固定索引,每次都只取第一行的内容。 - 遍历
rows时,没有针对当前行提取单元格,而是依赖全局的data变量。 - 获取最后陈述链接时,
LastStatementLinks = table.find_all("a", href=True)取的是整个表格的所有链接,再固定取LastStatementLinks[1],自然每次都是同一个链接。
修复代码
遍历行时,只提取当前行的单元格,且从当前行内获取链接:
# 移除全局的data = table.find_all('td') rows = table.find_all('tr') for row in rows[1:]: # 跳过表头行 cells = row.find_all('td') # 仅获取当前行的单元格 ExecutionID = str(cells[0].get_text()) Lastname = str(cells[3].get_text()) Firstname = str(cells[4].get_text()) TDJC = str(cells[5].get_text()) Age = str(cells[6].get_text()) Date = str(cells[7].get_text()) Race = str(cells[8].get_text()) County = str(cells[9].get_text()) # 从当前行获取最后陈述链接 last_statement_a = cells[2].find('a', href=True) if last_statement_a: LastStatementLink = last_statement_a.get("href") Urlcomplete = UrlLastStatement + LastStatementLink print(Urlcomplete) # ...后续获取陈述、插入数据库的代码
问题2:部分囚犯最后陈述链接错误
现象
部分囚犯(如545、544、552)的链接被错误拼接,例如囚犯545的链接变成https://www.tdcj.texas.gov/death_row/death_row/dr_info/cardenasrubenlast.html,导致无法访问。
原因
- 赋值错误:if语句里用了
==(比较运算符)而非=(赋值运算符),比如linkLS == "xxx"根本没修改变量值。 - 硬编码修复逻辑有缺陷:修改特殊ID的链接后,没有执行后续的陈述获取逻辑,导致这几个ID的陈述内容为空。
修复方案
方案1:统一处理链接(推荐)
使用urljoin自动处理路径重复问题,无需针对单个ID硬编码:
from urllib.parse import urljoin base_url = 'https://www.tdcj.texas.gov' # ... for row in table.find_all("tr")[1:]: cells = row.find_all('td') # ...其他字段提取 # 用urljoin拼接链接,自动处理重复前缀 linkLS = urljoin(base_url, cells[2].a['href'])
方案2:修正硬编码赋值错误
如果坚持用ID判断修复,要修正赋值符号,并且统一执行陈述获取逻辑:
for row in table.find_all("tr")[1:]: cells = row.find_all('td') ExecutionID= str(cells[0].get_text()) # ...其他字段提取 linkLS = 'https://www.tdcj.texas.gov/death_row/' + cells[2].a['href'] # 修正特殊ID的链接 if ExecutionID == "545": linkLS = "https://www.tdcj.texas.gov/death_row/dr_info/cardenasrubenlast.html" elif ExecutionID == "544": linkLS = "https://www.tdcj.texas.gov/death_row/dr_info/pruettrobertlast.html" elif ExecutionID == "552": linkLS = "https://www.tdcj.texas.gov/death_row/dr_info/no_last_statement.html" # 统一执行陈述获取逻辑 try: response = requests.get(linkLS, verify=False) print ("--- retrieving statement", ExecutionID," ---") print (linkLS) soup = BeautifulSoup(response.text, "html.parser") pattern = re.compile("Last Statement:") statement = soup.find(string=pattern).findNext('p').contents[0] # ...清理陈述内容的代码 except Exception: print("An error has occured while retrieving the statement. Proceeding...") statement = "" # ...插入数据库的代码
额外优化
- 移除冗余的
urllib.request.urlopen,直接用requests.get返回的响应文本解析。 - 修正表格创建语句的语法错误:
InmateInformation, text改为InmateInformation text(缺少空格)。
内容的提问来源于stack exchange,提问作者Omnicron
相关产品推荐
相关产品推荐

