Araara ca — Akurtosks
This is the sixth in a series of expository blog posts on Araara. The preceding post introduced alphabets as Asks, including gapped alphabets. The series begins here.
Normal form gives us an agreed way to represent each integer. But it can be quite long! It would be nice to have a shorter equivalent, especially when we are writing entire strings.
For example, the Unicode code point of ‘i’ is 105. Its Unicode normal form is:
uiffdcbb = −1094 + 1094 + 41 + 41 + 14 + 5 + 2 + 2 = 105
Could we reach the same value by a shorter route?

To find a shorter route, let’s split our alphabet into two parts.
Two gapped alphabets
Recall the Aglobasa alphabet:
o ae bp cj dt fv gk hx iu lr mn sz wy
We will keep all twenty-five positions, but form two gapped alphabets. The first keeps only u, replacing every other character with a gap. The second keeps every character except u, replacing it with a gap.
The letters retain their original positions and values. In particular, u still has value −1094.
Using the notation from the preceding post, the u alphabet is:
ui ui ui ui ui ui ui ui ui ui ui ui ui ui ui ui uiffddcb ui ui ui ui ui ui ui ui
And the non-u alphabet is:
uiffdda uiffda uiffdc uiffdb uiffddb uiffdba uiffdcc uiffdbb uiffddca uiffdca uiffddcba uiffdcb uiffdcca uiffdcba uiffddcc uiffdcbb ui uiffdccb uiffddbb uiffdccba uiffdd uiffddc uig uiffddcbb uiffddcca
Remember that ui indicates a gap. The term uiffddcb encodes the character ‘u’, whose Unicode code point is 117. That code point is distinct from the value −1094 assigned to u in Aglobasa.
These two alphabets have signatures too:
uiffddcb — u alphabet: 117
uiiihfdcbb — non-u alphabet: 2617
Their signatures add up to the signature of the full Aglobasa alphabet:
117 + 2617 = 2734
Expressed in Unicode normal form, the result is uiiihgfdbb, just as in the preceding post. Each character occurs in exactly one of our two gapped alphabets, and the gaps contribute zero.
A product of alphabets
We can now consider the Cartesian product of these two alphabets.
Think of choosing an ice cream: one choice of cone, and one choice of flavour. Together they make a pair.
A term of our product likewise has two components: a term using the u alphabet, and a term using the non-u alphabet. Their order matters.

We will use the notation:
u alphabet × non-u alphabet
For short Unicode notation, we adopt the convention that the first component is exactly one u.
The second component uses only non-u letters, retaining their Aglobasa positions and values. We will choose a shortest term in this alphabet.
So our representations will have a particularly simple shape: u, followed by a term containing no u.
Short Unicode notation
The initial u contributes −1094. To represent a code point N, the remaining term must therefore contribute N + 1094.
Here is how we choose that term:
- Find the fewest non-u letters that sum to N + 1094. Both positive and negative letters are allowed.
- Arrange each candidate in reverse Aglobasa alphabetical order.
- Among candidates of that length, choose the alphabetically first, comparing from left to right in Aglobasa order.
Then prepend u. We will call the result the short Unicode form.
For example, code point 13 has Unicode normal form uiccba. But:
uide = −1094 + 1094 + 14 − 1 = 13
And our opening example becomes:
uigtpe = −1094 + 1094 + 122 − 14 − 2 − 1 = 105
We could also use the suffix igtjb, which has the same length and value as igtpe. We choose igtpe because its first differing letter is p rather than j, and p comes first in Aglobasa.
The order here is Aglobasa’s, not English’s!
From Asks to Akurtosks
Recall that an Asks concatenates code points in Unicode normal form. Its shorter cousin will be an Akurtosks, formed by concatenating their short Unicode forms.
Here is ‘Unicode’ as an Asks:
uiffba uiffdd uiffdcbb uiffdba uiffdda uiffdbb uiffdc
And as an Akurtosks:
uiffba uigtb uigtpe uigttc uigjje uigvdc uiffdc
That takes 41 letters instead of 47. Both expressions encode the same sequence of code points, and both have signature 711.
The Aglobasa alphabet itself gives us a larger example. Its Asks representation in the preceding post takes 183 letters. As an Akurtosks, it becomes:
uigjje uiffda uiffdc uiffdb uigjj uigttc uigtp uigvdc uigje uigtje uigpp uigtj uigte uigtpp uigp uigtpe uigj uigt uigjpe uigta uigtb uigjp uig uigpe uige
That takes 129 letters: 54 fewer!
The spaces are just for readability. Each u starts a new encoded code point, because the second component contains no u letters. We still need no terminators.
An Asks uses terms of one alphabet. An Akurtosks instead uses terms of our product: each encoded code point combines the fixed u component with a shortest non-u component.
Same characters, same order, same sum. A shorter way to write them!
Exercises
- The character ‘u’ has code point 117. Check that its short Unicode form is uigj.
- Code point zero has Unicode normal form uio. What is its short Unicode form?
- Encode a short word as both an Asks and an Akurtosks. How many letters do you save? Check that their signatures agree.
In the next post, we will see some more things we can reference with Akurto.