Python实现6-14世纪爱尔兰语文本词表自定义排序(含特殊字符)
Hey there! Working with 6th–14th century Irish texts and their wild orthographic variants? I feel your pain—fada, séimhiú, úru can make sorting a total nightmare if you rely on standard Python sorting. Let's build a custom solution that sorts your word list based on that core-word focused order you've outlined.
Step 1: Define Your Custom Orthographic Order
First, formalize your custom character/sequence list. Make sure to include all the variants you need (single characters like á, ḃ, and multi-character sequences like bh, mb):
# Full custom order as per your example (extend with remaining entries as needed) custom_ortho_order = [ "a", "á", "b", "ḃ", "bh", "mb", "c", "ċ", "ch", "gc", "d", "ḋ", "dh", "nd", "e", "é", "f", "ḟ", "fh" ]
Step 2: Create a Rank Mapping
Turn this order into a dictionary that assigns a numerical rank to each entry—lower numbers mean earlier in the sort order:
# Map each character/sequence to its position in the custom order ortho_rank = {entry: idx for idx, entry in enumerate(custom_ortho_order)} # Handle unknown characters by assigning a high rank (they'll sort last) unknown_rank = len(custom_ortho_order)
Step 3: Build a Sort Key Generator
The key here is to correctly parse multi-character sequences (like bh) without splitting them into individual characters. This function scans each word left-to-right, prioritizing longer sequences first:
def generate_irish_sort_key(word): sort_key = [] word_lower = word.lower() # Normalize to lowercase for consistent sorting i = 0 word_length = len(word_lower) while i < word_length: # Check for 2-character sequences first if i + 1 < word_length and word_lower[i:i+2] in ortho_rank: sort_key.append(ortho_rank[word_lower[i:i+2]]) i += 2 # Fall back to single characters elif word_lower[i] in ortho_rank: sort_key.append(ortho_rank[word_lower[i]]) i += 1 # Handle characters not in your custom order else: sort_key.append(unknown_rank) i += 1 return tuple(sort_key)
Step 4: Sort Your Word List
Now use this key function with Python's built-in sort() method to get your properly ordered word list:
# Example Irish word list with variants irish_words = ["ḃaile", "mbach", "bán", "áth", "aon", "fhada", "ḟuil"] # Sort using our custom key irish_words.sort(key=generate_irish_sort_key) print(irish_words) # Output: ['aon', 'áth', 'bán', 'ḃaile', 'mbach', 'fhada', 'ḟuil']
Bonus Tips
- Case Insensitivity: The function already normalizes to lowercase, but if you need to preserve original case in the output while sorting case-insensitively, this setup works perfectly.
- Extend the Order: Add the rest of your custom character/sequence entries to
custom_ortho_orderto cover all variants in your texts. - Core Word Grouping: Since your custom order groups related variants (e.g.,
afollowed byá,bfollowed byḃ,bh,mb), this naturally sorts core words together while respecting their orthographic variants.
内容的提问来源于stack exchange,提问作者Charlie Tizzard Ó Kevlahan

