如何通过Python爬虫获取Futbin网站的球员个人图片链接?
Hey there! Let's fix that missing player image link issue in your scraping code.
First, let's break down what's happening: your current code targets all images starting with the Futbin CDN path, which grabs club, nationality, and league logos—but the player profile image is either in a different position in that list, or better yet, we can target it directly with a more precise selector to avoid confusion.
Modified Code with Player Image Extraction
Here's the updated version of your code that pulls the player's profile image, along with your existing image grabs:
import requests import bs4 for page in range(1, 5): req = requests.get('https://www.futbin.com/players?page='+str(page)) soup = bs4.BeautifulSoup(req.text,'html.parser') numbers = [1,2] for number in numbers: players = soup.findAll('tr','player_tr_'+str(number)) for p in players: # Grab club, nationality, league images (your original code, with safety check) img_elements = p.select('img[src^="https://cdn.futbin.com/content/fifa22/img/"]') if len(img_elements) >= 3: clubImage = img_elements[0]['src'] nationalityImage = img_elements[1]['src'] leagueImage = img_elements[2]['src'] # Grab the player's profile image player_img_tag = p.select_one('td.player_img img') if player_img_tag: # Handle lazy-loaded images (some sites use data-src instead of src) player_image_link = player_img_tag.get('data-src') or player_img_tag['src'] print(player_image_link) # Test output to verify
Key Changes Explained
- Precise Selector:
td.player_img imgtargets the specific<td>with classplayer_img(where Futbin stores the player avatar) and its child<img>tag. This is more reliable than relying on index positions, which can break if the page structure changes. - Lazy Load Handling: Some sites use
data-srcfor lazy-loaded images instead ofsrc. Theget('data-src') or player_img_tag['src']line checks for both, so you'll get the correct URL regardless of lazy loading. - Safety Checks: Added
if len(img_elements) >=3andif player_img_tagto prevent errors if elements are missing (e.g., if the page structure updates unexpectedly).
Quick Note
Always make sure to respect Futbin's robots.txt and terms of service—don't send too many requests too quickly, or you might get your IP blocked. Consider adding delays between requests using time.sleep() if needed.
内容的提问来源于stack exchange,提问作者BarCode

