Python新手求助:HTML文件人名关联链接提取函数无返回值问题
Hey there! Let's work through this step by step since you're just starting out with Python—no worries about the code being rough, that's totally normal when you're learning! Let's fix up your function to extract those name-linked URLs and add a proper return value.
First, let's iron out the basic issues in your current code, then build out the logic to pull the links you need and return them.
Step 1: Fix Basic Code Problems
- Move your library imports to the top of your script (this is standard Python practice, rather than importing inside the function).
- Use a
withstatement to open the file—it automatically closes the file when you're done, which is safer and cleaner. - You had a typo: your function parameter is
file, but you usedfilenameinside the function—let's fix that mismatch. - Complete the
bs(f.read())call (you had a trailing...).
Step 2: Add Logic to Extract Name-Associated Links
The exact way to target these links depends on your HTML structure. For example, maybe names are in <a> tags directly, or wrapped in a <span> with a class like person-name. Here's a flexible example you can adjust to match your webpage:
Full Working Example
# Import required libraries at the top of your script import codecs from bs4 import BeautifulSoup as bs def get_links(file_path): # Use a with statement to safely open and read the HTML file with codecs.open(file_path, 'r', 'utf-8') as f: # Parse the HTML content with BeautifulSoup soup = bs(f.read(), 'html.parser') # 'html.parser' is built-in, no extra installs needed # Adjust this selector to match your actual HTML! # Example 1: If names are directly in <a> tags (narrow down if needed) # name_links = soup.find_all('a') # Example 2: If names are in <a> tags inside elements with class "person-name" name_links = soup.select('.person-name a') # Extract the name and URL from each matching link name_url_pairs = [] for link in name_links: # Get the clean name text (strip removes extra spaces/newlines) person_name = link.get_text(strip=True) # Get the URL from the href attribute person_url = link.get('href') # Only add to our list if both name and URL exist if person_name and person_url: name_url_pairs.append({ 'name': person_name, 'url': person_url }) # Return the list of name-url pairs as the function's output return name_url_pairs # Test the function with your file file = 'C:/users/me/webpage.html' results = get_links(file) # Print out the results to check for item in results: print(f"Name: {item['name']} | URL: {item['url']}")
Key Tips for You:
- Tweak the selector: Use your browser's inspect tool to look at the HTML around the name links. If names are in a different structure (like
<div class="profile"><a href="...">Alice</a></div>), update thesoup.select()orsoup.find_all()call to match. - Relative URLs: If the
hrefvalues are relative (like/alice-profile), you can convert them to absolute URLs by adding the base webpage URL (e.g.,base_url = "https://example.com"thenfull_url = base_url + person_url). - Return value: The function now returns a list of dictionaries, which makes it easy to use the data later (like saving to a CSV or processing further).
Remember, as a beginner, writing "imperfect" code is part of the process—you're already doing great by building something that solves a problem for you!
内容的提问来源于stack exchange,提问作者Miguel 2488

