如何调整爬取www.oddsportal.com的Python爬虫代码以处理赛事无比分场景
Got it, let's resolve that annoying data shifting issue when matches don't have a score yet. The root problem is we're hardcoding TD element indices, but rows without a table-score cell have one fewer column—so our odds end up pulling from the wrong positions. Here's how to adjust the logic properly:
Step 1: Modify the generate_matches Function
We need to dynamically check if a score exists in each row before pulling data, instead of relying on fixed indices. Update the function like this:
def generate_matches(table): tr_tags = table.findAll('tr') country = "" league = "" for tr_tag in tr_tags: if 'class' in tr_tag.attrs and 'dark' in tr_tag['class']: # This is a league header row, update country/league th_tag = tr_tag.find('th', {'class': 'first2 tl'}) a_tags = th_tag.findAll('a') country = a_tags[0].text league = a_tags[1].text else: td_tags = tr_tag.findAll('td') if len(td_tags) < 4: # Skip empty/malformed rows continue # Check if this row has a score cell has_score = any('table-score' in td.get('class', []) for td in td_tags) # Assign values based on whether score exists time = td_tags[0].text.strip() game = td_tags[1].text.strip() if has_score: score = td_tags[2].text.strip() home_odds = td_tags[3].text.strip() draw_odds = td_tags[4].text.strip() away_odds = td_tags[5].text.strip() else: score = None # Will convert to NaN in DataFrame home_odds = td_tags[2].text.strip() draw_odds = td_tags[3].text.strip() away_odds = td_tags[4].text.strip() yield time, game, score, home_odds, draw_odds, away_odds, country, league
Step 2: Keep the parse_data Loop Clean
Since we've already handled the score check in generate_matches, you don't need to change anything in the loop where you populate GameData—it can stay exactly as it was:
for row in generate_matches(table): game_data.date.append(game_date) game_data.time.append(row[0]) game_data.game.append(row[1]) game_data.score.append(row[2]) game_data.home_odds.append(row[3]) game_data.draw_odds.append(row[4]) game_data.away_odds.append(row[5]) game_data.country.append(row[6]) game_data.league.append(row[7])
What Changed?
- We added a check for the
table-scoreclass in each row's TD elements to detect if a score exists. - If no score is found, we set
scoretoNone(which Pandas automatically converts toNaNin the DataFrame) and shift the odds indices left by one. - We also added a check for empty rows (
len(td_tags) <4) to avoid errors from malformed rows. - Added
.strip()to clean up any extra whitespace in the text values for cleaner data.
This way, whether a match has a score or not, your home_odds, draw_odds, and away_odds will always pull from the correct positions, and the score column will show NaN for matches without results.
内容的提问来源于stack exchange,提问作者user16304089

