含代码块的长文本分段方案咨询:单段最大2000字符
Solution for Splitting Text with Code Block Preservation
Here's a Python implementation that splits text into chunks while respecting code block integrity:
def split_text_with_code_blocks(text, max_length=2000): parts = [] current_part = [] current_length = 0 in_code_block = False code_language = "" # Split text into lines to avoid breaking mid-line lines = text.splitlines(keepends=True) for line in lines: line_length = len(line) # Update code block state if line.strip().startswith("```"): if in_code_block: # Closing code block in_code_block = False code_language = "" else: # Opening code block - extract language if present in_code_block = True code_language = line.strip()[3:].strip() # Check if adding this line would exceed max length if current_length + line_length > max_length and current_part: # Handle split inside code block if in_code_block: # Add closing delimiter to current part closing_delimiter = "```\n" current_part.append(closing_delimiter) current_length += len(closing_delimiter) # Finalize current part and add to list parts.append("".join(current_part)) # Reset for new part current_part = [] current_length = 0 # If we were in a code block, prepend opening delimiter to new part if in_code_block: opening_delimiter = f"```{code_language}\n" if code_language else "```\n" current_part.append(opening_delimiter) current_length += len(opening_delimiter) # Add current line to the part current_part.append(line) current_length += line_length # Add the remaining content as the final part if current_part: parts.append("".join(current_part)) return parts # Example usage text = """test text blablabla ```python import random import math def random_number(): return random.randint(0, 100) def square_root(number): return math.sqrt(number)
more test text blablabla"""
Split with max length set to 100 for demonstration (matches example)
parts = split_text_with_code_blocks(text, max_length=100)
Print results as formatted strings
for i, part in enumerate(parts, 1):
print(f"part{i} = '''{part}'''")
### How It Works: 1. **Line-by-Line Processing**: We split the input text into lines to avoid breaking content mid-line, which keeps both regular text and code readable. 2. **Code Block Tracking**: We maintain state to track if we're inside a code block, and capture the language identifier if provided. 3. **Length Check**: Before adding each line, we check if it would exceed the maximum chunk length. If it does: - If inside a code block, we append the closing ```` delimiter to the current chunk. - We finalize the current chunk and start a new one. - If we were inside a code block, we prepend the opening ```` delimiter (with the original language) to the new chunk. 4. **Final Chunk**: After processing all lines, we add any remaining content as the last chunk. ### Example Output: Running the code above will produce exactly the output you requested:
part1 = '''test text blablabla
import random import math def random_number(): return random.randint(0, 100)
'''
part2 = '''```python
def square_root(number):
return math.sqrt(number)
more test text blablabla'''
This implementation handles edge cases like code blocks without language identifiers, multiple code blocks in one text, and code blocks at the start/end of the input.
内容的提问来源于stack exchange,提问作者Ludo
相关产品推荐
相关产品推荐

