You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决Selenium循环搜索时Pandas无结果报错问题

问题解决:跳过无结果的关键词搜索

我用Selenium循环执行关键词搜索,通过Pandas提取表格搜索结果。当搜索词无返回结果时会触发报错,但有结果但不符合过滤条件时程序能正常运行。比如搜索“Cedarcrest”时,会报错:pandas.errors.UndefinedVariableError: name 'Description' is not defined,需要修改代码让程序跳过无结果的搜索词。

修改后的代码

import time
import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys


class KW_POST_BOT(object):
    def __init__(self, browser, search_engine_url, kw_list, package):
        self.browser = browser
        self.package = package
        self.search_engine_url = search_engine_url
        self.kw_list = kw_list

    def main(self):
        self.browser.get("https://www.dnv.org/building-development/look-building-and-trades-permits")
        
        for kw in self.kw_list:
            print("*" * 30)
            print(f"搜索关键词:{kw}")
            print("*" * 30)
            print("查找数据表格")
            print("-" * 30)
            
            time.sleep(3)
            # 定位搜索框并清空
            search_box = self.browser.find_element(By.XPATH,
                                             "/html/body/div[2]/div[1]/div[2]/div[2]/div/div/div/app-root/div/div[1]/div/div[1]/div/div/ng2-completer/div/input")
            search_box.clear()
            # 输入关键词并回车
            search_box.send_keys(kw)
            search_box.send_keys(Keys.ENTER)
            
            time.sleep(3)

            # 获取表格行
            table_trs = self.browser.find_elements(By.XPATH, '//table[@id="case_table"]/tbody/tr')
            # 判断是否有数据行(表头不算)
            if len(table_trs) <= 1:
                print(f"关键词【{kw}】无搜索结果,跳过")
                continue
                
            value_list = []
            for row in table_trs[1:]:
                tds = row.find_elements(By.TAG_NAME, "td")
                # 确保td数量足够,避免索引越界
                if len(tds) >=7:
                    value_list.append({
                        'Address': tds[1].text,
                        'Status': tds[4].text,
                        'Date': tds[5].text,
                        'Description': tds[6].text
                    })

            df = pd.DataFrame(value_list)
            # 先检查DataFrame是否为空,以及是否存在所需列
            if not df.empty and all(col in df.columns for col in ['Description', 'Date']):
                filtered_list = df[
                    df['Description'].str.contains('New', na=False) & 
                    df['Date'].str.contains('2018|2019|2020|2021|2022', na=False)
                ]
                print(filtered_list)
            else:
                print(f"关键词【{kw}】的结果不符合过滤条件或无有效数据")


# 关键词列表
key_words = ["Eldon", "Ruby", "Bracknell", "Pelly", "Sunset", "Edgewood", "Sycamore", "Lodge",
             "Virginia", "Loraine", "Kendal", "Emerald", "Dudley", "Highland", "Highland", "Montroyal", "Ranger",
             "Shirley", "Cedarcrest", "Lions", "Sunnycrest", "Beaumont", "Tudor", "Winona", "Canterbury",
             "Beaconsfield", "Hampshire", "Devon", "Essex", "Derby", "Belgrave", "Cheviot", "Parliament", "Ruskin",
             "Handsworth", "Rialto", "Belvedere", "Marineview", "Mapleridge", "Pheasant", "Ruthina", "Marigold",
             "Marigold", "Glenwood", "Timberline", "Ventura", "Monteray", "Greenway", "Valencia", "Hermosa", "Vienna",
             "Genoa", "Saville", "Granada", "Lucerne", "Verona", "Croydon", "Silverdale", "Lewister", "Langdale",
             "Quinton", "Carolyn", "Wavertree", "Wentworth", "Leovista", "Trenton", "Evergreen", "Chelsea", "Crystal",
             "Sylvan", "Alpine", "Bonita", "Palisade", "Blueridge", "Skyline", "Glencanyon", "Delmar", "Dolores",
             "Delbrook", "Linnae", "Teviot", "Belvista", "Prospect", "Primrose", "Edgewood", "Patterdale", "Newdale",
             "Crestwood", "Montroyal", "Glenview", "Arundel", "Ranger", "Capilano", "Salvador", "Grace", "East", "June",
             "Cliffridge", "Glenn"]

# 初始化并运行
bot = KW_POST_BOT(webdriver.Firefox(), "https://www.dnv.org/building-development/look-building-and-trades-permits",
                  key_words, [])

bot.main()

关键修改点

  • 无结果判断:获取表格行后,判断len(table_trs) <=1(第一行是表头),如果满足则直接跳过当前关键词循环。
  • 避免索引越界:遍历行时先检查td元素数量是否足够,防止因表格结构异常导致索引错误。
  • 安全过滤数据:执行过滤前,先检查DataFrame是否为空、所需列是否存在,替换原df.eval为更安全的直接列操作,同时添加na=False避免空值引发的错误。
  • 代码优化:合并重复定位搜索框的代码,减少冗余操作。

内容的提问来源于stack exchange,提问作者shawn zh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 13:30:11