Python循环迭代链接爬取网页时遇TypeError问题求助
问题解决
错误原因
你遇到的 TypeError: 'int' object is not iterable 是因为尝试遍历单个整数 torneo = 0,而整数并非可迭代对象。你的目标是遍历多个整数(0-3)获取对应赛事链接,因此需要将整数替换为可迭代的列表。
修正后的代码
from typing import Generator from requests_html import HTMLSession import pandas as pd import numpy as np from itertools import islice import mysql.connector from sqlalchemy import create_engine def _get_rows(url: str) -> Generator[dict[str, str], None, None]: session = HTMLSession() # 使用传入的url参数,而非全局变量 r = session.get(url) allmatch = r.html.find(".in-match") results = r.html.find(".h-text-center a") matchodds = r.html.find("[data-odd]") odds = [matchodd.attrs["data-odd"] for matchodd in matchodds] idx = 0 N = 20 for match, res in islice(zip(allmatch, results), N): if res.text in ["POSTP.", "0:2 ABN.", "0:3 AWA.", "3:0 AWA.", "0:2 CAN.", "2:0 CAN.", "1:1 CAN."]: continue print(f"{match.text} Z {res.text} {', '.join(odds[idx:idx+3])}") yield { "match": match.text, "result": res.text, "odds": ", ".join(odds[idx : idx + 3]), "best_bets": odds[idx], "oddtwo": odds[idx+1], "oddthree": odds[idx+2], } idx += 3 if __name__ == "__main__": # 定义需要遍历的赛事ID列表 tournament_ids = [0, 1, 2, 3] # 可选:收集所有赛事数据到一个DataFrame all_dfs = [] for torneo in tournament_ids: # 根据赛事ID获取对应链接和标识 match torneo: case 0: matchlinkok = "https://www.betexplorer.com/football/chile/primera-division/results/" camp = "Chi-A" case 1: matchlinkok = "https://www.betexplorer.com/football/algeria/ligue-1/results/" camp = "Alg-A" case 2: matchlinkok = "https://www.betexplorer.com/football/australia/a-league/results/" camp = "Aus-A" case 3: matchlinkok = "https://www.betexplorer.com/football/austria/bundesliga/results/" camp = "Aut-A" case _: print("无效的赛事ID") continue matchlink = matchlinkok NM = 20 # 爬取数据并生成DataFrame df = pd.DataFrame(_get_rows(matchlink)) df[['home','away']] = df['match'].str.split("-", expand=True) df[['scorehome','scoreaway']] = df['result'].str.split(":", expand=True) df = df.astype({'scorehome':'int', 'scoreaway':'int'}) # 添加赛事相关统计字段 cols = ['scorehome', 'scoreaway'] df['tournament'] = camp df['sum_stats'] = df[cols].sum(axis=1, numeric_only=True) df['over05'] = np.where(df['sum_stats']>0, 'OK', 'NO') df['over15'] = np.where(df['sum_stats']>1, 'OK', 'NO') df['over25'] = np.where(df['sum_stats']>2, 'OK', 'NO') df['over35'] = np.where(df['sum_stats']>3, 'OK', 'NO') df['over45'] = np.where(df['sum_stats']>4, 'OK', 'NO') df['goal'] = np.where((df['scorehome'] >= 1) & (df['scoreaway'] >= 1), 'OK', 'NO') df['esito'] = [ '1' if h > a else '2' if h < a else 'X' for h, a in zip(df['scorehome'], df['scoreaway'])] # 格式化结果字段并生成唯一标识 df['result'] = df['result'].str.replace(':','-') colss = ['home', 'away', 'result'] df['uniquefield'] = df[colss].apply(lambda row: ' '.join(row.values.astype(str)), axis=1) df = df[['home','away','scorehome','scoreaway','result','best_bets','oddtwo','oddthree', 'sum_stats', 'over05', 'over15', 'over25', 'over35', 'over45', 'goal', 'esito', 'tournament', 'uniquefield']] print(f"\n===== 赛事 {camp} 数据 =====") print(df) all_dfs.append(df) # 可选:合并所有赛事数据 combined_df = pd.concat(all_dfs, ignore_index=True) print("\n===== 所有赛事合并数据 =====") print(combined_df)
关键修改点
- 修复遍历问题:将单个整数
torneo=0替换为可迭代的赛事ID列表tournament_ids = [0,1,2,3],通过for torneo in tournament_ids遍历每个赛事。 - 匹配整数类型:将
match语句中的字符串案例(如"0")改为整数案例(如0),与遍历的ID类型一致。 - 修复函数参数问题:
_get_rows函数原本忽略传入的url参数,改为使用session.get(url)确保每个赛事使用对应的链接。 - 优化条件判断:将多个重复的
if res.text == ...合并为一个in语句,简化代码。 - 循环内处理数据:将所有数据处理逻辑移入循环内部,确保每个赛事都能独立爬取和处理。
内容的提问来源于stack exchange,提问作者Franky77
相关产品推荐
相关产品推荐

