You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python循环迭代链接爬取网页时遇TypeError问题求助

问题解决

错误原因

你遇到的 TypeError: 'int' object is not iterable 是因为尝试遍历单个整数 torneo = 0,而整数并非可迭代对象。你的目标是遍历多个整数(0-3)获取对应赛事链接,因此需要将整数替换为可迭代的列表。

修正后的代码

from typing import Generator
from requests_html import HTMLSession
import pandas as pd
import numpy as np
from itertools import islice
import mysql.connector
from sqlalchemy import create_engine

def _get_rows(url: str) -> Generator[dict[str, str], None, None]:
    session = HTMLSession()
    # 使用传入的url参数,而非全局变量
    r = session.get(url)

    allmatch = r.html.find(".in-match")
    results = r.html.find(".h-text-center a")
    matchodds = r.html.find("[data-odd]")

    odds = [matchodd.attrs["data-odd"] for matchodd in matchodds]

    idx = 0
    N = 20
    for match, res in islice(zip(allmatch, results), N):

        if res.text in ["POSTP.", "0:2 ABN.", "0:3 AWA.", "3:0 AWA.", "0:2 CAN.", "2:0 CAN.", "1:1 CAN."]:
            continue

        print(f"{match.text} Z {res.text} {', '.join(odds[idx:idx+3])}")
        yield {
            "match": match.text,
            "result": res.text,
            "odds": ", ".join(odds[idx : idx + 3]),
            "best_bets": odds[idx], 
            "oddtwo": odds[idx+1], 
            "oddthree": odds[idx+2],
        }

        idx += 3

if __name__ == "__main__":
    # 定义需要遍历的赛事ID列表
    tournament_ids = [0, 1, 2, 3]
    # 可选:收集所有赛事数据到一个DataFrame
    all_dfs = []

    for torneo in tournament_ids:
        # 根据赛事ID获取对应链接和标识
        match torneo:   
            case 0:
                matchlinkok = "https://www.betexplorer.com/football/chile/primera-division/results/"
                camp = "Chi-A"
            case 1:
                matchlinkok = "https://www.betexplorer.com/football/algeria/ligue-1/results/"
                camp = "Alg-A"
            case 2:
                matchlinkok = "https://www.betexplorer.com/football/australia/a-league/results/"
                camp = "Aus-A" 
            case 3:
                matchlinkok = "https://www.betexplorer.com/football/austria/bundesliga/results/"
                camp = "Aut-A"   
            case _:
                print("无效的赛事ID")
                continue
        
        matchlink = matchlinkok
        NM = 20

        # 爬取数据并生成DataFrame
        df = pd.DataFrame(_get_rows(matchlink))
        df[['home','away']] = df['match'].str.split("-", expand=True)     
        df[['scorehome','scoreaway']] = df['result'].str.split(":", expand=True) 
        df = df.astype({'scorehome':'int', 'scoreaway':'int'})
        
        # 添加赛事相关统计字段
        cols = ['scorehome', 'scoreaway']
        df['tournament'] = camp
        df['sum_stats'] = df[cols].sum(axis=1, numeric_only=True)
        df['over05'] = np.where(df['sum_stats']>0, 'OK', 'NO')
        df['over15'] = np.where(df['sum_stats']>1, 'OK', 'NO')
        df['over25'] = np.where(df['sum_stats']>2, 'OK', 'NO')
        df['over35'] = np.where(df['sum_stats']>3, 'OK', 'NO')
        df['over45'] = np.where(df['sum_stats']>4, 'OK', 'NO')
        df['goal'] = np.where((df['scorehome'] >= 1) & (df['scoreaway'] >= 1), 'OK', 'NO')
        df['esito'] = [ '1' if h > a else '2' if h < a else 'X' for h, a in zip(df['scorehome'], df['scoreaway'])]
        
        # 格式化结果字段并生成唯一标识
        df['result'] = df['result'].str.replace(':','-')
        colss = ['home', 'away', 'result']
        df['uniquefield'] = df[colss].apply(lambda row: ' '.join(row.values.astype(str)), axis=1)
        df = df[['home','away','scorehome','scoreaway','result','best_bets','oddtwo','oddthree', 'sum_stats', 'over05', 'over15', 'over25', 'over35', 'over45', 'goal', 'esito', 'tournament', 'uniquefield']]

        print(f"\n===== 赛事 {camp} 数据 =====")
        print(df)
        all_dfs.append(df)
    
    # 可选:合并所有赛事数据
    combined_df = pd.concat(all_dfs, ignore_index=True)
    print("\n===== 所有赛事合并数据 =====")
    print(combined_df)

关键修改点

  1. 修复遍历问题:将单个整数 torneo=0 替换为可迭代的赛事ID列表 tournament_ids = [0,1,2,3],通过 for torneo in tournament_ids 遍历每个赛事。
  2. 匹配整数类型:将 match 语句中的字符串案例(如 "0")改为整数案例(如 0),与遍历的ID类型一致。
  3. 修复函数参数问题:_get_rows 函数原本忽略传入的 url 参数,改为使用 session.get(url) 确保每个赛事使用对应的链接。
  4. 优化条件判断:将多个重复的 if res.text == ... 合并为一个 in 语句,简化代码。
  5. 循环内处理数据:将所有数据处理逻辑移入循环内部,确保每个赛事都能独立爬取和处理。

内容的提问来源于stack exchange,提问作者Franky77

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 22:27:17