You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用lxml.builder将URL以列表形式展示并生成HTML?

Solution: Generate HTML List from URLs with lxml.builder

Got it, let's work through this to turn your string of YouTube URLs into a clean HTML list using lxml's builder API. Here's a complete, functional version of your main() function that does exactly what you need:

def main():
    from lxml.builder import E
    import re

    p_persons = []
    
    # Your sample person data
    person = [
        '1',  # Counter
        'ID',
        '0. https://www.youtube.com/watch?v=qLsn5aNaVkI 1. https://www.youtube.com/watch?v=MPbO6P3Vtx8 2. https://www.youtube.com/watch?v=jVKWPaFuNng 3. https://www.youtube.com/watch?v=9HFyB4gCOqY 4. https://www.youtube.com/watch?v=muQGef4Df_8',
        'birthplace'
    ]
    
    # Step 1: Extract all valid YouTube URLs from the string
    url_string = person[2]
    # Regex grabs every URL that follows "number. " pattern
    urls = re.findall(r'\d+\. (https://\S+)', url_string)
    
    # Step 2: Build an HTML ordered list (ol) of clickable links
    url_list = E.ol(
        *[E.li(E.a(f'Video {i+1}', href=url)) for i, url in enumerate(urls)]
    )
    
    # Step 3: Assemble the full person entry with all details
    person_entry = E.div(
        E.p(f'Counter: {person[0]}'),
        E.p(f'ID: {person[1]}'),
        E.p('Linked Videos:'),
        url_list,
        E.p(f'Birthplace: {person[3]}')
    )
    
    p_persons.append(person_entry)
    
    # Optional: Convert the lxml elements to a readable HTML string
    from lxml import etree
    print(etree.tostring(person_entry, pretty_print=True, encoding='unicode'))

if __name__ == "__main__":
    main()

Breakdown of Key Steps:

  • URL Extraction: The regex r'\d+\. (https://\S+)' targets every URL prefixed with a number and dot, so we only pull valid links from your messy string.
  • List Building: Using a list comprehension, we generate an <li> for each URL, with an <a> tag that links directly to the video. Wrapping in <ol> keeps the numbering matching your original data.
  • Structured HTML: We wrap all person details in a <div> for clean organization—you can swap out <p> tags for headers like <h4> if you want more visual hierarchy.
  • Preview Output: The etree.tostring() call converts the lxml elements into a human-readable HTML string, which you can save to a file or integrate into a larger page.

If you prefer an unordered list instead, just replace E.ol() with E.ul()—everything else works the same.

内容的提问来源于stack exchange,提问作者Suncatcher

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:42:38