Araara ḃ — Asks

Share
Araara ḃ — Asks

This is the fourth in a series of expository blog posts on Araara. The preceding post introduced the Aglobasa alphabet and the Akurto type that uses it. The series begins here.

Alphabets are arbitrary

Before proceeding, I should make one convention explicit. Choosing Aglobasa as our default alphabet is a convention. Its sounds make it convenient to read aloud, but the arithmetic does not depend on that particular choice.

o ae bp cj dt fv gk hx iu lr mn sz wy

But that’s kind of the point! I might prefer Aglobasa, you might prefer English, and someone else might want Greek letters. Once we specify the (ordered!) alphabet and the rule assigning its values, we can interpret each other’s numerals.

The type name Akurto is simply a translation of Ashort into Globasa. By convention, I use Aglobasa for Akurto and ø-English for Ashort. We already know how to write integers in normal form. Let’s use that to represent Unicode code points, then whole strings—and, in the next post, alphabets themselves.

A woman in a peach cardigan writes with a pen on paper at a wooden desk, against a soft pastel-blue wash.
Alphabets are arbitrary.

Unicode normal form

For a Unicode code point N, define its Unicode normal form (unf) by writing ui followed by the ordinary Akurto normal form of N.

For example, uppercase ‘A’ has Unicode code point 65. Its Akurto normal form is:

fdcc = 41 + 14 + 5 + 5 = 65

Adding the prefix gives its Unicode normal form:

uifdcc = −1094 + 1094 + 41 + 14 + 5 + 5 = 65

The prefix ui has value −1094 + 1094 = 0. It gives the numeral a recognisable beginning while preserving its value.

To encode a string:

  1. For a Unicode code point N:
    1. Write N in ordinary Akurto normal form.
    2. Prepend ui. Its value is zero.
    3. To form a string, concatenate the resulting Unicode normal forms in order.

This gives each code point one agreed representation, using the normal form we already know.

Inherent boundaries

The prefix also tells us where each encoded code point begins. Unicode code points are nonnegative, so their normal forms contain positive letters, or o for zero. They never contain u, and therefore never contain ui.

Each occurrence of ui starts a new numeral. Its suffix continues until the next ui, or the end of the string. The boundaries are inherent, even without spaces.

Here are three examples:

0 = uio (unf)

11 = uicca (unf)

13 = uiccba (unf)

Four rounded pastel tiles displaying A, Greek omega, a heart, and a star on ivory paper.
Some Unicode characters

Akurto Unicode

A Unicode code point is an integer between zero and 1,114,111. It is usually written in hexadecimal with a ‘U+’ prefix. For example, uppercase ‘A’ has code point U+0041, or 65 in decimal. See the Unicode definition of a code point and the Basic Latin chart.

A visible character can consist of several code points; we encode those in order. Our opening example, ‘A’, needs just one:

‘A’ = uifdcc

A few characters

The first three uppercase English letters have code points 65, 66, and 67. Their Unicode normal forms are:

‘A’ = uifdcc

‘B’ = uifdcca

‘C’ = uifdccb

Integers

Here is something cool. In the first post, we set out to represent the integers in Araara. To do this, we defined Ashort, and then Akurto. Now consider

uillihhffdb

This points to U+2124, which is ℤ! In fact, metonymically, uillihhffdb is a name for the integers themselves! We use the close association between the double-struck Z in Unicode and the integers to refer to the latter by naming the former.

Forming strings

We can now form strings by concatenating code points in Unicode normal form.

For example, ‘Unicode’ has code points 85, 110, 105, 99, 111, 100, and 101. Encoding them one at a time gives:

‘Unicode’ = uiffba uiffdd uiffdcbb uiffdba uiffdda uiffdbb uiffdc

The spaces are just for readability. Remove them and each ui still identifies the next code point. Concatenation preserves their order, so we can recover the original sequence.

Even though Akurto has only twenty-five letters, we can express any sequence of Unicode code points. Isn’t that cool? We will call these Araara character code sequences (Accs) in English and Araara simbol-kodi silsila in Globasa (Asks).

Zero is a code point too

Code point zero has normal form o, so its Unicode normal form is uio. That final o represents zero; it is part of the numeral, not a terminator.

The prefix ui alone is not a complete Unicode normal form: a code point still needs its normal-form suffix. For example, two zero code points concatenate as uiouio.

What if we sum an Asks?

As in the preceding posts, we can also evaluate the whole concatenated expression as one integer. For ‘Unicode’, the code points sum to 711:

uiffba uiffdd uiffdcbb uiffdba uiffdda uiffdbb uiffdc

= hggffdca = 365 + 122 + 122 + 41 + 41 + 14 + 5 + 1 = 711

We could consider hggffdca a simple numerical signature, or a very basic hash, of ‘Unicode’. Given the string, its sum follows unambiguously. Each ui contributes zero, so evaluating the concatenation gives exactly the sum of the code points.

The reverse does not hold. Every anagram has the same sum, and even strings that are not anagrams can collide: ‘AC’ and ‘BB’ both sum to 132. The sum preserves neither order nor the individual characters. Still, it is a neat consequence of using the same arithmetic for numbers and text.

We will call such sums the (additive) signature of an Asks.

An elegant handwritten signature in charcoal ink, with flowing loops over pale blue and peach washes.
A signature

Exercises

  1. ‘B’ has code point 66. Find its Unicode normal form, then encode the string ‘AB’ by concatenating the two numerals.
  2. Compare the encodings and sums of ‘AB’ and ‘BA’. What information survives concatenation, and what disappears when we take the sum?

We can now write integers, code points, and strings using Akurto. That brings us one step closer to describing the alphabets themselves in the next post.