You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于HTML模板生成格式化纯文本文件(Django/Python实现)

从模板生成纯文本文件的Python实现方案(含批量处理)

方案一:直接使用Jinja2编写纯文本模板

这是最高效的方案——跳过HTML模板环节,直接针对纯文本格式编写Jinja2模板,完全匹配输出需求,无需额外格式转换。

纯文本模板示例(book_template.txt)

{{ book.name }}
The author of the book is {{ book.author }}.
For more description, go to "{{ book.url }}"
  Some important notions on the book
    {% for item in points %}
    - {{ item }}
    {% endfor %}

Python代码实现(含批量生成)

from jinja2 import Environment, FileSystemLoader
import os

# 初始化Jinja2环境,指定模板目录
env = Environment(loader=FileSystemLoader('templates'))
template = env.get_template('book_template.txt')

# 模拟从数据库批量获取的数据集(实际项目中用ORM查询)
books_data = [
    {
        "book": {"name": "Python For Beginners", "author": "John Doe", "url": "https://mybooks.example.com/book1"},
        "points": ["Variable", "Strings", "File R/W operations", "Loops", "Conditions"]
    },
    # 更多书籍数据...
]

# 创建输出目录
os.makedirs('output_books', exist_ok=True)

# 批量生成文件
for idx, context in enumerate(books_data, 1):
    text_content = template.render(context)
    with open(f'output_books/book_{idx}.txt', 'w', encoding='utf-8') as f:
        f.write(text_content)

方案二:先渲染HTML模板,再转换为纯文本

如果已有现成的HTML模板,无需重新编写纯文本模板,可先渲染HTML再转成纯文本格式。常用工具包括html2text或BeautifulSoup。

使用html2text实现

from jinja2 import Environment, FileSystemLoader
import html2text
import os

# 加载HTML模板
env = Environment(loader=FileSystemLoader('templates'))
html_template = env.get_template('book_template.html')

# 批量数据
books_data = [
    {
        "book": {"name": "Python For Beginners", "author": "John Doe", "url": "https://mybooks.example.com/book1"},
        "points": ["Variable", "Strings", "File R/W operations", "Loops", "Conditions"]
    },
    # 更多数据...
]

# 初始化html2text转换器,配置格式
text_converter = html2text.HTML2Text()
text_converter.body_width = 0  # 禁用自动换行
text_converter.ignore_images = True
text_converter.ignore_links = False  # 保留链接文本

os.makedirs('output_books_html', exist_ok=True)

# 批量生成
for idx, context in enumerate(books_data, 1):
    html_content = html_template.render(context)
    text_content = text_converter.handle(html_content)
    # 手动修正列表格式,匹配需求
    text_content = text_content.replace('  * ', '    - ')
    with open(f'output_books_html/book_{idx}.txt', 'w', encoding='utf-8') as f:
        f.write(text_content)

使用BeautifulSoup自定义转换

如果需要更精细的格式控制,可直接解析HTML标签生成纯文本:

from jinja2 import Environment, FileSystemLoader
from bs4 import BeautifulSoup
import os

env = Environment(loader=FileSystemLoader('templates'))
html_template = env.get_template('book_template.html')

def html_to_text(html):
    soup = BeautifulSoup(html, 'html.parser')
    lines = []
    for elem in soup.recursiveChildGenerator():
        if isinstance(elem, str):
            text = elem.strip()
            if text:
                lines.append(text)
        elif elem.name == 'li':
            lines.append(f'    - {elem.get_text(strip=True)}')
        elif elem.name == 'p':
            lines.append(elem.get_text(strip=True))
    return '\n'.join(lines)

# 批量生成逻辑同前
for idx, context in enumerate(books_data, 1):
    html_content = html_template.render(context)
    text_content = html_to_text(html_content)
    with open(f'output_books_custom/book_{idx}.txt', 'w', encoding='utf-8') as f:
        f.write(text_content)

方案三:Django框架下的实现

Django自带模板系统支持纯文本渲染,也可复用现有HTML模板转文本。

直接渲染纯文本模板

from django.template import loader
from myapp.models import Book
import os

# 批量获取数据库数据(用prefetch_related避免N+1查询)
books = Book.objects.prefetch_related('notions').all()

# 加载纯文本模板
template = loader.get_template('book_text_template.txt')

os.makedirs('django_output', exist_ok=True)

for book in books:
    context = {
        'book': book,
        'points': [notion.content for notion in book.notions.all()]
    }
    text_content = template.render(context)
    with open(f'django_output/{book.slug}.txt', 'w', encoding='utf-8') as f:
        f.write(text_content)

Django中HTML转纯文本

可复用html2text或Django自带的strip_tags函数(后者仅去除标签,格式需手动调整):

from django.template import loader
from django.utils.html import strip_tags
from myapp.models import Book

template = loader.get_template('book_html_template.html')
books = Book.objects.all()

for book in books:
    html_content = template.render({'book': book, 'points': book.notions.all()})
    raw_text = strip_tags(html_content)
    # 手动调整格式,匹配输出需求
    formatted_text = raw_text.replace('\n\n', '\n').replace('Some important notions on the book', '\n  Some important notions on the book\n    - ')
    with open(f'django_output/{book.slug}.txt', 'w', encoding='utf-8') as f:
        f.write(formatted_text)

批量生成的优化建议

  1. 数据库查询优化:使用批量查询(如Django的all()、prefetch_related)避免频繁数据库请求,解决N+1查询问题。
  2. 并行处理:针对超100份文件的生成,可使用concurrent.futures.ThreadPoolExecutor(IO密集型任务适合线程)提升速度:
    from concurrent.futures import ThreadPoolExecutor
    
    def generate_file(context, idx):
        text_content = template.render(context)
        with open(f'output_books/book_{idx}.txt', 'w', encoding='utf-8') as f:
            f.write(text_content)
    
    with ThreadPoolExecutor(max_workers=4) as executor:
        executor.map(generate_file, books_data, range(1, len(books_data)+1))
    
  3. 异步任务:如果是Web应用,可使用Celery将批量生成任务放入异步队列,避免阻塞用户请求。

内容的提问来源于stack exchange,提问作者Steve Yonkeu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 15:10:41