迁移bash curl脚本到Python解析HTML提取门店名遇IndexError如何解决
解决方案
方法1:调用curl+grep+sed管道实现
可以通过subprocess的shell模式直接执行拼接好的管道命令,示例代码如下:
import subprocess store = "你的门店名称" # 直接拼接对应shell命令 shell_cmd = f"curl -silent https://example.com/location/{store} | grep -P -o 'Location.*?<div class=\"row\">' | sed -e :a -e 's/<[^>]*>//g;/</N;//ba' | sed 's/./& - /12'" # 执行命令捕获输出 res = subprocess.run(shell_cmd, shell=True, capture_output=True, text=True, encoding="utf-8") locations = res.stdout.strip() print(locations)
注意:如果门店名称包含特殊字符,需要提前做转义处理,避免shell注入风险。
方法2:修复BeautifulSoup解析逻辑(更推荐,无外部依赖)
报错原因
你遇到的IndexError: list index out of range有两个诱因:
- 代码存在笔误:你定义的会话变量是
session,但发起请求时调用了未定义的s变量 - 部分
<td>标签没有子内容,直接取contents[0]时因列表为空触发越界;同时直接遍历所有<td>没有按行分组,无法精准匹配对应列的内容
修正后代码
import requests from bs4 import BeautifulSoup store = 'Here Store Name' url = f"https://example.com/?store={store}" session = requests.Session() # 修正会话变量调用错误 resp = session.get(url) soup = BeautifulSoup(resp.text, "html.parser") store_names = [] # 按行遍历表格内容 for tr in soup.find_all("tr"): # 提取当前行所有td标签 tds = tr.find_all("td") # 过滤不足4列的行(表头、空行都自动过滤) if len(tds) < 4: continue # 提取第二列(下标从0开始计数,所以是1)的文本,自动去除空白字符 current_store_name = tds[1].text.strip() store_names.append(current_store_name) # 输出所有提取到的门店名 print(store_names)
内容的提问来源于stack exchange,提问作者joop
相关产品推荐
相关产品推荐

