Colab中Selenium文本提取结果异常问题求助
Oddsportal爬虫跨环境结果异常问题
我使用Python的Chromedriver爬取指定赛事页面时出现以下问题:
- 在Spyder本地环境运行脚本一切正常
- 在Colab中执行时,Selenium能定位到目标元素,但提取出网站不存在的异常数值
- 在Replit中运行时,网站整体加载内容出现错误,代码本身无问题
测试与核心函数
元素存在性检测函数fi
def fi(a): try: driver.find_element("xpath", a).text except: return False
文本提取函数ffi
def ffi(a): if fi(a) != False : return driver.find_element("xpath", a).text
完整爬虫代码
driver.get("https://www.oddsportal.com/soccer/croatia/hnl/hnk-gorica-varazdin-Kr4sLgwt/#1X2;2") for j in range(1,15): print(j) book= ffi('((//*[starts-with(@class,"flex text-xs max")])[{}]//p)[1]'.format(j)) if fi('((//*[starts-with(@class,"flex text-xs max")])[{}]//p)[2]//preceding-sibling::a'.format(j))==False: Odd_1=ffi('((//*[starts-with(@class,"flex text-xs max")])[{}]//p)[2]'.format(j)) else: Odd_1=fi('((//*[starts-with(@class,"flex text-xs max")])[{}]//a)[5]'.format(j)) if fi('((//*[starts-with(@class,"flex text-xs max")])[{}]//p)[3]//preceding-sibling::a'.format(j))==False: Odd_X=ffi('((//*[starts-with(@class,"flex text-xs max")])[{}]//p)[3]'.format(j)) else: Odd_X=ffi('((//*[starts-with(@class,"flex text-xs max")])[{}]//a)[6]'.format(j)) if fi('((//*[starts-with(@class,"flex text-xs max")])[{}]//p)[4]//preceding-sibling::a'.format(j))==False: Odd_2=ffi('((//*[starts-with(@class,"flex text-xs max")])[{}]//p)[4]'.format(j)) else: Odd_2=ffi('((//*[starts-with(@class,"flex text-xs max")])[{}]//a)[7]'.format(j)) ab= (ffi('//div[contains(@class,"flex items-center w-full h-auto")]//p')) bc=(ffi('(//div[contains(@class,"flex px")]//child::div)[3]')) print(book, Odd_1, Odd_X, Odd_2,ab ,bc)
结果对比说明
- 期望结果:显示正常的博彩公司名称、胜平负赔率,以及赛事基本信息
- Colab运行结果:提取出的赔率等数值为网站不存在的异常值,出现不符合逻辑的数字
- Replit运行结果:页面加载内容错误,无法显示正常的赛事赔率列表,整体页面结构混乱
内容的提问来源于stack exchange,提问作者Benjamin Fuchs
相关产品推荐
相关产品推荐

