You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取HTML中的链接?附示例代码求助

Hey there! Extracting links from HTML with Python is a super common task, and using a dedicated HTML parsing library like BeautifulSoup is way more reliable than trying to hack it with regex (regex for HTML is a messy rabbit hole—trust me, I've been there). Let's break down how to solve your specific problem step by step.

First, let's recap the HTML snippet you shared:

<td><img src="/images/cflags/png/id1.png" alt="Indonesia" title="Indonesia"></td> <td></td> <td>link.here/python.php </td> <td>Linux</td> <td><img src="/images/cflags/png/id2.png" alt="Indonesia" title="Indonesia"></td> <td></td> <td>link2.here/python.php </td> <td>Linux</td> <td><img src="/images/cflags/png/id3.png" alt="Indonesia" title="Indonesia"></td> <td></td> <td>link3.here/python.php </td> <td>Linux</td>

Step 1: Install BeautifulSoup

First, you'll need to install the beautifulsoup4 library if you haven't already. Run this command in your terminal:

pip install beautifulsoup4

Step 2: Write the Python Code

Here's a complete script to extract those links. We'll parse the HTML, find all <td> elements, and filter out the ones that contain your target links:

from bs4 import BeautifulSoup

# Your HTML content as a string
html_content = """
<td><img src="/images/cflags/png/id1.png" alt="Indonesia" title="Indonesia"></td> <td></td> <td>link.here/python.php </td> <td>Linux</td> <td><img src="/images/cflags/png/id2.png" alt="Indonesia" title="Indonesia"></td> <td></td> <td>link2.here/python.php </td> <td>Linux</td> <td><img src="/images/cflags/png/id3.png" alt="Indonesia" title="Indonesia"></td> <td></td> <td>link3.here/python.php </td> <td>Linux</td>
"""

# Parse the HTML using BeautifulSoup
soup = BeautifulSoup(html_content, 'html.parser')

# Extract links: filter <td> elements whose text ends with .php, and clean up whitespace
links = [td.get_text(strip=True) for td in soup.find_all('td') if td.get_text(strip=True).endswith('.php')]

# Print the results
print("Extracted links:")
for link in links:
    print(link)

What This Does:

  • BeautifulSoup(html_content, 'html.parser') parses the raw HTML string into a navigable object that's easy to work with.
  • soup.find_all('td') grabs every <td> element from the HTML snippet.
  • The list comprehension filters out only the <td> elements whose text (after stripping extra whitespace) ends with .php—perfect for targeting your specific links.
  • strip=True removes any extra spaces around the link text (like the trailing space in your original snippet).

Alternative: Using Fixed Positions

If your HTML structure is always consistent (links are every 4th <td>, starting at the 3rd one), you could also grab them by index:

all_td_elements = soup.find_all('td')
# Links are at positions 2, 6, 10 (0-indexed)
links = [all_td_elements[i].get_text(strip=True) for i in [2, 6, 10]]

This works if you're 100% sure the structure never changes, but the first method is more flexible if the HTML layout might vary over time.


内容的提问来源于stack exchange,提问作者Taher

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:15:07