如何用BeautifulSoup提取网站特定子类?提取球员姓名求助
Hey there! Let's tackle your two questions step by step.
1. Extracting Specific Child Elements with BeautifulSoup
BeautifulSoup gives you several straightforward ways to target specific child elements:
.find()/.find_all(): Use these to search for elements by tag name, class, ID, or other attributes. For example, to get all<p>tags inside a<div>with classcontent:content_div = soup.find('div', class_='content') specific_paragraphs = content_div.find_all('p', class_='highlight')- CSS Selectors with
.select(): This is great for more complex hierarchies. For instance, to get the second<li>child of a<ul>with IDmenu:target_li = soup.select_one('ul#menu li:nth-child(2)') - Direct Child Traversal: Use
.childrenor.findChildren(recursive=False)(like you did in your code) to get only immediate children of an element, ignoring nested ones.
2. Fixing Your Code to Extract Only Player Names
Looking at your current code, you're skipping the first column entirely with [1:], but we need to dig into that first column to pull out just the player's name. Here's how to adjust it:
First, let's modify the loop to handle the first column separately, since it likely contains extra elements (like images or links) alongside the name. Assuming the player name is inside an <a> tag (common on soccer market value sites), here's the updated code:
import csv import requests from bs4 import BeautifulSoup def few(urls, file): f = open(file, 'a', newline='', encoding="utf-8") writer = csv.writer(f) url = urls page = requests.get(url, headers={'User-Agent': 'Mozilla/5.0'}) soup = BeautifulSoup(page.content, 'lxml') tbody = soup('table', {"class": "items"})[0].find_all('tr') for row in tbody: # Skip header rows if needed (uncomment if your table has a header tr) # if row.find('th'): # continue # Extract player name from the first column first_col = row.findChildren(recursive=False)[0] # Adjust the selector below based on the actual HTML structure # If the name is in a span with class 'player-name', use: # player_name = first_col.find('span', class_='player-name').text.strip() player_name = first_col.find('a').text.strip() # Process remaining columns as before, excluding unwanted indices other_cols = row.findChildren(recursive=False)[1:] exclude = [0, 3, 4, 6, 7, 8] filtered_cols = [ele.text.strip() for idx, ele in enumerate(other_cols) if idx not in exclude] # Combine name and filtered columns, then write to CSV writer.writerow([player_name] + filtered_cols) f.close() # Don't forget to close the file!
Key Notes:
- Adjust the selector for player name: If the name isn't in an
<a>tag, right-click the player name on the website → "Inspect" to find its exact HTML element (e.g., a<span>with a specific class). Then update thefirst_col.find(...)line to match. - Skip header rows: If your table has a header row (with
<th>tags), uncomment the check to skip it so you don't write header text into your CSV. - Close the file: Added
f.close()to ensure your data is properly saved.
内容的提问来源于stack exchange,提问作者Gofos
相关产品推荐
相关产品推荐

