You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现6-14世纪爱尔兰语文本词表自定义排序(含特殊字符)

Custom Sorting for Irish Historical Text Analysis in Python

Hey there! Working with 6th–14th century Irish texts and their wild orthographic variants? I feel your pain—fada, séimhiú, úru can make sorting a total nightmare if you rely on standard Python sorting. Let's build a custom solution that sorts your word list based on that core-word focused order you've outlined.

Step 1: Define Your Custom Orthographic Order

First, formalize your custom character/sequence list. Make sure to include all the variants you need (single characters like á, ḃ, and multi-character sequences like bh, mb):

# Full custom order as per your example (extend with remaining entries as needed)
custom_ortho_order = [
    "a", "á", "b", "ḃ", "bh", "mb",
    "c", "ċ", "ch", "gc", "d", "ḋ", "dh", "nd",
    "e", "é", "f", "ḟ", "fh"
]

Step 2: Create a Rank Mapping

Turn this order into a dictionary that assigns a numerical rank to each entry—lower numbers mean earlier in the sort order:

# Map each character/sequence to its position in the custom order
ortho_rank = {entry: idx for idx, entry in enumerate(custom_ortho_order)}
# Handle unknown characters by assigning a high rank (they'll sort last)
unknown_rank = len(custom_ortho_order)

Step 3: Build a Sort Key Generator

The key here is to correctly parse multi-character sequences (like bh) without splitting them into individual characters. This function scans each word left-to-right, prioritizing longer sequences first:

def generate_irish_sort_key(word):
    sort_key = []
    word_lower = word.lower()  # Normalize to lowercase for consistent sorting
    i = 0
    word_length = len(word_lower)
    
    while i < word_length:
        # Check for 2-character sequences first
        if i + 1 < word_length and word_lower[i:i+2] in ortho_rank:
            sort_key.append(ortho_rank[word_lower[i:i+2]])
            i += 2
        # Fall back to single characters
        elif word_lower[i] in ortho_rank:
            sort_key.append(ortho_rank[word_lower[i]])
            i += 1
        # Handle characters not in your custom order
        else:
            sort_key.append(unknown_rank)
            i += 1
    return tuple(sort_key)

Step 4: Sort Your Word List

Now use this key function with Python's built-in sort() method to get your properly ordered word list:

# Example Irish word list with variants
irish_words = ["ḃaile", "mbach", "bán", "áth", "aon", "fhada", "ḟuil"]

# Sort using our custom key
irish_words.sort(key=generate_irish_sort_key)

print(irish_words)
# Output: ['aon', 'áth', 'bán', 'ḃaile', 'mbach', 'fhada', 'ḟuil']

Bonus Tips

  • Case Insensitivity: The function already normalizes to lowercase, but if you need to preserve original case in the output while sorting case-insensitively, this setup works perfectly.
  • Extend the Order: Add the rest of your custom character/sequence entries to custom_ortho_order to cover all variants in your texts.
  • Core Word Grouping: Since your custom order groups related variants (e.g., a followed by á, b followed by ḃ, bh, mb), this naturally sorts core words together while respecting their orthographic variants.

内容的提问来源于stack exchange,提问作者Charlie Tizzard Ó Kevlahan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:17:34