You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决ChatGPT3.5 Turbo生成表格时的Token超限问题

解决方案

方案1:直接解析为结构化数据生成表格(推荐)

如果表格文本是空格/制表符分隔的规范格式,完全不需要调用LLM,用Python的pandas直接解析并生成Markdown表格,彻底规避token限制问题,且数据零丢失。

示例代码:

import pandas as pd
from io import StringIO

# 从UI获取的原始表格文本
raw_table_text = """bird_id bird_posts bird_likes
012 2 5
034 10 22
056 7 15
078 15 30
..."""  # 可支持任意大体积数据

# 读取为DataFrame(自动识别空格分隔的列与表头)
df = pd.read_csv(StringIO(raw_table_text), sep=r'\s+', header=0)

# 生成标准Markdown表格
markdown_table = df.to_markdown(index=False)
print(markdown_table)

方案2:分块处理表格文本(适配非规范格式)

如果表格格式混乱(比如存在合并单元格、不规则分隔)必须依赖LLM解析,可按行分块,每块保留表头,确保LLM能正确识别列映射,最后合并结果避免数据丢失。

示例代码:

import tiktoken
from openai import OpenAI

client = OpenAI(api_key="你的API密钥")

def count_tokens(text: str) -> int:
    # 计算文本对应的token数
    encoder = tiktoken.encoding_for_model("gpt-3.5-turbo")
    return len(encoder.encode(text))

def split_table_into_chunks(table_text: str, max_token_limit: int = 900) -> list[str]:
    # 拆分表格文本,每块保留表头,预留指令token空间
    lines = [line.strip() for line in table_text.split('\n') if line.strip()]
    if not lines:
        return []
    
    header = lines[0]
    chunks = []
    current_chunk = [header]
    current_token_count = count_tokens(header)

    for line in lines[1:]:
        line_token_count = count_tokens(line)
        # 预留100token给指令文本,避免触发模型上限
        if current_token_count + line_token_count + 100 <= max_token_limit:
            current_chunk.append(line)
            current_token_count += line_token_count
        else:
            chunks.append('\n'.join(current_chunk))
            current_chunk = [header, line]
            current_token_count = count_tokens(header) + line_token_count
    
    # 添加最后一块数据
    if current_chunk:
        chunks.append('\n'.join(current_chunk))
    return chunks

def process_chunk_with_llm(chunk: str) -> str:
    # 调用LLM处理单块表格文本
    response = client.chat.completions.create(
        model="gpt-3.5-turbo",
        messages=[{"role": "user", "content": f"Create a Markdown table with the given text:\n{chunk}"}],
        temperature=0  # 固定输出格式,避免偏差
    )
    return response.choices[0].message.content.strip()

# 主执行流程
raw_table_text = "你的大体积原始表格文本"
table_chunks = split_table_into_chunks(raw_table_text)
processed_chunks = [process_chunk_with_llm(chunk) for chunk in table_chunks]

# 合并分块结果:仅保留第一个表头,跳过后续块的重复表头与分隔线
final_table_lines = []
for idx, chunk_table in enumerate(processed_chunks):
    chunk_lines = chunk_table.split('\n')
    if idx == 0:
        final_table_lines.extend(chunk_lines)
    else:
        # 跳过Markdown表格的表头行和分隔线行(前两行)
        final_table_lines.extend(chunk_lines[2:])

final_markdown_table = '\n'.join(final_table_lines)
print(final_markdown_table)

方案3:使用大token上限模型

直接切换到支持更大token容量的模型(如gpt-3.5-turbo-16k,支持16384token),无需分块预处理,直接处理大体积表格文本。

示例代码:

from openai import OpenAI

client = OpenAI(api_key="你的API密钥")

response = client.chat.completions.create(
    model="gpt-3.5-turbo-16k",  # 替换为大token模型
    messages=[{"role": "user", "content": f"Create a Markdown table with the given text:\n{你的大体积表格文本}"}],
    temperature=0
)

print(response.choices[0].message.content)

方案选型建议

  • 优先选方案1:速度快、无API成本、数据准确率100%,适合规范格式的表格。
  • 选方案2:仅当表格格式极不规范,必须依赖LLM语义解析时使用,保证行列信息不丢失。
  • 选方案3:适合快速实现,且可接受稍高API成本的场景。

内容的提问来源于stack exchange,提问作者usr_lal123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 06:43:17