You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

迁移bash curl脚本到Python解析HTML提取门店名遇IndexError如何解决

解决方案

方法1:调用curl+grep+sed管道实现

可以通过subprocess的shell模式直接执行拼接好的管道命令,示例代码如下:

import subprocess

store = "你的门店名称"
# 直接拼接对应shell命令
shell_cmd = f"curl -silent https://example.com/location/{store} | grep -P -o 'Location.*?<div class=\"row\">' | sed -e :a -e 's/<[^>]*>//g;/</N;//ba' | sed 's/./& - /12'"
# 执行命令捕获输出
res = subprocess.run(shell_cmd, shell=True, capture_output=True, text=True, encoding="utf-8")
locations = res.stdout.strip()
print(locations)

注意:如果门店名称包含特殊字符,需要提前做转义处理,避免shell注入风险。

方法2:修复BeautifulSoup解析逻辑(更推荐,无外部依赖)

报错原因

你遇到的IndexError: list index out of range有两个诱因:

  1. 代码存在笔误:你定义的会话变量是session,但发起请求时调用了未定义的s变量
  2. 部分<td>标签没有子内容,直接取contents[0]时因列表为空触发越界;同时直接遍历所有<td>没有按行分组,无法精准匹配对应列的内容

修正后代码

import requests
from bs4 import BeautifulSoup

store = 'Here Store Name'
url = f"https://example.com/?store={store}"
session = requests.Session()
# 修正会话变量调用错误
resp = session.get(url)
soup = BeautifulSoup(resp.text, "html.parser")

store_names = []
# 按行遍历表格内容
for tr in soup.find_all("tr"):
    # 提取当前行所有td标签
    tds = tr.find_all("td")
    # 过滤不足4列的行(表头、空行都自动过滤)
    if len(tds) < 4:
        continue
    # 提取第二列(下标从0开始计数,所以是1)的文本,自动去除空白字符
    current_store_name = tds[1].text.strip()
    store_names.append(current_store_name)

# 输出所有提取到的门店名
print(store_names)

内容的提问来源于stack exchange,提问作者joop

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 21:24:01